← Quick search of all subject terms
MICROBIOLOGY · REFERENCE DESK

Microbiologynoun explanation · Page 5

202bilingual terms · Current number 5 / 7The

This collection combines attributed Wikipedia excerpts and original SciAtlas bilingual definitions under CC BY-SA 4.0. Excerpts were extracted and shortened; machine-assisted Chinese translations are labeled. Original entries provide further reading. Language versions may differ in emphasis and do not replace standards. Concepts can appear in several disciplines; consult standards and original literature for rigorous use.

contains 202 terms · This page displays 30 terms, you can enter keywords to query the complete range
Microbiology

FASTA format

FASTA格式

在生物信息学中,FASTA格式是一种用于记录核酸序列或肽序列的文本格式,其中的核酸或氨基酸均以单个字母编码呈现。该格式同时还允许在序列之前定义名称和编写注释。这一格式最初由FASTA软件包定义,但现今已是生物信息学领域的一项标准。 FASTA简明的格式降低了序列操纵和分析的难度,令序列可被文本处理工具和诸如Python、Ruby和Perl等脚本语言处理。

In bioinformatics and biochemistry, the FASTA format is a text-based format for representing either nucleotide sequences or amino acid (protein) sequences, in which nucleotides or amino acids are represented using single-letter codes. The format allows for sequence names and comments to precede the sequences. It originated from the FASTA software package and has since become a near-universal standard in bioinformatics. The simplicity of FASTA format makes it easy to manipulate and parse sequences using text-processing tools and scripting languages.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements.

View content license ↗
Microbiology

FASTQ format

FASTQ格式

FASTQ格式是一种保存生物序列(通常为核酸序列)及其测序质量得分信息的文本格式。序列与质量得分皆由单个ASCII字符表示。 该格式最初由维尔康姆基金会桑格研究所开发,旨在将FASTA格式序列及其质量数据集成在一起。而目前,FASTQ格式已经成为了保存高通量测序结果的事实标准。

FASTQ format is a text-based format for storing both a biological sequence (usually nucleotide sequence) and its corresponding quality scores. Both the sequence letter and quality score are each encoded with a single ASCII character for brevity. It was originally developed at the Wellcome Trust Sanger Institute to bundle a FASTA formatted sequence and its quality data, but has become the de facto standard for storing the output of high-throughput sequencing instruments such as the Illumina Genome Analyzer.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements.

View content license ↗
Microbiology

GC-content

GC含量

GC含量(GC-content,guanine-cytosine content)是分子生物学和遗传学的术语,指研究对象(例如放线菌)的全基因组(DNA 或 RNA 分子)或其片段中,含氮碱基鸟嘌呤(G)或胞嘧啶(C)任何一个所占的百分比。一种生物的基因组或特定DNA、RNA片段有特定的GC含量。 在DNA链中G和C是以三个氢键相连,而T和A则是两个氢键相连的。氢键的多少体现连接的能量,氢键多的不容易被打断。 在双链DNA中,腺嘌呤与胸腺嘧啶(A/T)之比,以及鸟嘌呤与胞嘧啶(G/C)之比都是1。但是,(A+T)/(G+C)之比则随DNA的种类不同而异。GC含量愈高,DNA的密度也愈高,同时热及碱不易使之变性,因此利用这一特性便可进行DNA的分离或测定。 测定GC含量的方法有:Tm法,HPLC法

In molecular biology and genetics, GC-content (or G+C content or guanine-cytosine content) is the percentage of nitrogenous bases in a DNA or RNA molecule that are either guanine (G) or cytosine (C). This measure indicates the proportion of G and C bases out of an implied four total bases, also including adenine and thymine in DNA and adenine and uracil in RNA. GC-content may be given for a certain fragment of DNA or RNA or for an entire genome. When it refers to a fragment, it may denote the GC-content of an individual gene or section of a gene (domain), a group of genes or gene clusters, a non-coding region, or a synthetic oligonucleotide such as a primer.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements.

View content license ↗
Microbiology

GC skew

气相色谱偏差

GC 偏差是指 DNA 或 RNA 的特定区域中核苷酸鸟嘌呤和胞嘧啶过多或不足。 GC 偏差也是一种测量链特异性鸟嘌呤过度表达的统计方法。在平衡条件下(没有突变或选择压力,并且核苷酸随机分布在基因组内),DNA 分子的两条单链上的四种 DNA 碱基(腺嘌呤、鸟嘌呤、胸腺嘧啶和胞嘧啶)的频率相等。然而,在大多数细菌(例如大肠杆菌)和一些古细菌(例如硫磺菌)中,前导链和滞后链之间的核苷酸组成是不对称的:前导链含有更多的鸟嘌呤(G)和胸腺嘧啶(T),而滞后链含有更多的腺嘌呤(A)和胞嘧啶(C)。

GC skew is when the nucleotides guanine and cytosine are over- or under-abundant in a particular region of DNA or RNA. GC skew is also a statistical method for measuring strand-specific guanine overrepresentation. In equilibrium conditions (without mutational or selective pressure and with nucleotides randomly distributed within the genome) there is an equal frequency of the four DNA bases (adenine, guanine, thymine, and cytosine) on both single strands of a DNA molecule. However, in most bacteria (e.g. E. coli) and some archaea (e.g. Sulfolobus solfataricus), nucleotide compositions are asymmetric between the leading strand and the lagging strand: the leading strand contains more guanine (G) and thymine (T), whereas the lagging strand contains more adenine (A) and cytosine (C).

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

K-mer

K一聚体

在生物信息学中,k-mers 是生物序列中包含的长度为 k {\displaystyle k} 的子串。主要用于计算基因组学和序列分析,其中 k 聚体由核苷酸(即 A、T、G 和 C)组成,k 聚体可用于组装 DNA 序列、改善异源基因表达、识别宏基因组样本中的物种以及创建减毒疫苗。通常,术语 k-mer 是指长度为 k {\displaystyle k} 的序列的所有子序列,因此序列 AGAT 将具有四个单体(A、G、A 和 T)、三个 2-mer(AG、GA、AT)、两个 3-mer(AGA 和 GAT)和一个 4-mer(AGAT)。

In bioinformatics, k-mers are substrings of length k {\displaystyle k} contained within a biological sequence. Primarily used within the context of computational genomics and sequence analysis, in which k-mers are composed of nucleotides (i.e. A, T, G, and C), k-mers are capitalized upon to assemble DNA sequences, improve heterologous gene expression, identify species in metagenomic samples, and create attenuated vaccines. Usually, the term k-mer refers to all of a sequence's subsequences of length k {\displaystyle k} , such that the sequence AGAT would have four monomers (A, G, A, and T), three 2-mers (AG, GA, AT), two 3-mers (AGA and GAT) and one 4-mer (AGAT).

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Base calling

碱基检出

碱基识别是将核碱基分配给色谱峰、光强度信号或核苷酸通过纳米孔引起的电流变化的过程。完成这项工作的计算机程序是 Phred,它是学术和商业 DNA 测序实验室广泛使用的碱基识别软件程序,因为它具有很高的碱基识别准确性。目前,碱基检出通常由仪器上的软件处理,例如专有的实时分析 (RTA) 管道,该管道高度集成并随每个平台版本进行更新。纳米孔测序的碱基识别器(例如 Guppy 或 Dorado)使用根据从准确测序数据获得的当前信号进行训练的神经网络。

Base calling is the process of assigning nucleobases to chromatogram peaks, light intensity signals, or electrical current changes resulting from nucleotides passing through a nanopore. One computer program for accomplishing this job is Phred, which was a widely used base calling software program by both academic and commercial DNA sequencing laboratories because of its high base calling accuracy. Currently basecalling is commonly handled by on-instrument software, such as the proprietary Real-Time Analysis (RTA) pipeline, which is highly integrated and updated with each platform release. Base callers for Nanopore sequencing like Guppy or Dorado, use neural networks trained on current signals obtained from accurate sequencing data.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Contig

重叠群

重叠群(来自连续的)是一组重叠的 DNA 片段,它们一起代表 DNA 的共有区域。在自下而上的测序项目中,重叠群是指重叠的序列数据(reads);在自上而下的测序项目中,重叠群是指形成基因组物理图谱的重叠克隆,用于指导测序和组装。因此,重叠群可以指重叠的 DNA 序列,也可以指克隆中包含的重叠的物理片段(片段),具体取决于上下文。

A contig (from contiguous) is a set of overlapping DNA segments that together represent a consensus region of DNA. In bottom-up sequencing projects, a contig refers to overlapping sequence data (reads); in top-down sequencing projects, contig refers to the overlapping clones that form a physical map of the genome that is used to guide sequencing and assembly. Contigs can thus refer both to overlapping DNA sequences and to overlapping physical segments (fragments) contained in clones depending on the context.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Indel

插入缺失

Indel(插入-删除)是一个分子生物学术语,指生物体基因组中碱基的插入或删除。长度≥ 50 个碱基的插入缺失被归类为结构变体。在基因组的编码区,除非插入缺失的长度是3的倍数,否则就会产生移码突变。例如,导致移码的常见微插入会导致犹太人或日本人群中的布卢姆综合症。插入缺失可以与点突变进行对比。插入或删除是从序列中插入或删除核苷酸,而点突变是一种替换形式,替换其中一个核苷酸而不改变 DNA 中的总数。 Indels 也可以与串联碱基突变 (TBM) 进行对比,后者可能是由根本不同的机制引起的。

Indel (insertion-deletion) is a molecular biology term for an insertion or deletion of bases in the genome of an organism. Indels ≥ 50 bases in length are classified as structural variants. In coding regions of the genome, unless the length of an indel is a multiple of 3, it will produce a frameshift mutation. For example, a common microindel which results in a frameshift causes Bloom syndrome in the Jewish or Japanese population. Indels can be contrasted with a point mutation. An indel inserts or deletes nucleotides from a sequence, while a point mutation is a form of substitution that replaces one of the nucleotides without changing the overall number in the DNA. Indels can also be contrasted with Tandem Base Mutations (TBM), which may result from fundamentally different mechanisms.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Copy number variation

拷贝数变异

拷贝数变异(CNV)是一种基因组部分重复的现象,并且基因组中的重复次数因个体而异。拷贝数变异是一种结构变异:具体来说,它是一种影响大量碱基对的重复或删除事件。整个人类基因组的大约三分之二可能由重复组成,并且人类基因组的 4.8-9.5% 可归类为拷贝数变异。在哺乳动物中,拷贝数变异在群体和疾病表型中产生必要的变异方面发挥着重要作用。拷贝数变异通常可分为两大类:短重复和长重复。然而,两组之间没有明确的界限,分类取决于感兴趣基因座的性质。

Copy number variation (CNV) is a phenomenon in which sections of the genome are repeated and the number of repeats in the genome varies between individuals. Copy number variation is a type of structural variation: specifically, it is a type of duplication or deletion event that affects a considerable number of base pairs. Approximately two-thirds of the entire human genome may be composed of repeats and 4.8–9.5% of the human genome can be classified as copy number variations. In mammals, copy number variations play an important role in generating necessary variation in the population as well as disease phenotype. Copy number variations can be generally categorized into two main groups: short repeats and long repeats. However, there are no clear boundaries between the two groups and the classification depends on the nature of the loci of interest.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Structural variation

結構變異

基因组结构变异是生物体染色体结构的变异,例如缺失、重复、拷贝数变异、插入、倒位和易位。最初,结构变异影响的序列长度约为 1kb 至 3Mb,该长度大于 SNP,小于染色体异常(尽管定义有一些重叠)。然而,结构变体的操作范围已扩大到包括 > 50bp 的事件。一些结构变异与遗传疾病有关,但大多数则不然。大约13%的人类基因组在正常人群中被定义为结构变异,并且在人群中至少有240个基因以纯合缺失多态性存在,表明这些基因在人类中是可有可无的。

Genomic structural variation is the variation in structure of an organism's chromosome, such as deletions, duplications, copy-number variants, insertions, inversions and translocations. Originally, a structure variation affects a sequence length about 1kb to 3Mb, which is larger than SNPs and smaller than chromosome abnormality (though the definitions have some overlap). However, the operational range of structural variants has widened to include events > 50bp. Some structural variants are associated with genetic diseases, however most are not. Approximately 13% of the human genome is defined as structurally variant in the normal population, and there are at least 240 genes that exist as homozygous deletion polymorphisms in human populations, suggesting these genes are dispensable in humans.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Gene expression

基因表現

基因表达(英语:gene expression)又称基因表现,是用基因中的信息来合成基因产物的过程。产物通常是蛋白质,但对于非蛋白质编码基因,如tRNA和小核RNA(snRNA),产物则是RNA。所有已知生物都通过基因表达来生成生命所需的高分子物质。 基因表达的过程可概分为:DNA转录、RNA剪接、RNA转译、蛋白质转译后修饰,这四大步骤。基因表达调控控制细胞的结构与功能,同时也是细胞分化、形态发生及生物体的多功能性和适应性的基础。不同的时间、不同的环境,以及不同部位的细胞,或是基因在细胞中的含量差异,皆可能使基因产生不同的表现。基因调节也可以作为进化变化的底物,因为基因表达的时间,位置和数量的控制可以对基因在细胞或多细胞生物体中的功能(作用)具有深远的影响。 在遗传学中,基因表达是基因型产生表现型(即可观察的性状)的最基本的层次。

Gene expression is the process by which the information contained within a gene is used to produce a functional gene product, such as a protein or a functional RNA molecule. This process involves multiple steps, including the transcription of the gene's sequence into RNA. For protein-coding genes, this RNA is further translated into a chain of amino acids that folds into a protein, while for non-coding genes, the resulting RNA itself serves a functional role in the cell. Gene expression enables cells to utilize the genetic information in genes to carry out a wide range of biological functions. While expression levels can be regulated in response to cellular needs and environmental changes, some genes are expressed continuously with little variation.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements.

View content license ↗
Microbiology

Gene prediction

基因预测

基因预测(英语:gene prediction)或称基因发现(gene finding),是生物信息学的一个重要分支,使用生物学实验或计算机等手段识别DNA序列上的具有生物学特征的片段。基因识别的对象主要是蛋白质编码基因,也包括其他具有一定生物学功能的因子,如RNA基因和调控因子。基因识别是基因组研究的基础。 在早期,基因识别的主要手段是基于活的细胞或生物的实验。通过对若干种不同基因的同源重组的速率的统计分析,我们能够获知它们在染色体上的顺序。若进行大量类似的分析,我们可以确定各个基因的大致位置。现在,由于人类已经获得了巨大数量的基因组信息,依靠较慢的实验分析已不能满足基因识别的需要,而基于计算机算法的基因识别得到了长足的发展,成为了基因识别的主要手段。 识别具有生物学功能的片段与判定该片段(或其对应的产品)的功能是两个不同的概念,后者通常需要通过基因敲除等的实验手段来决定。不过,生物信息学的前沿研究正在使得由基因序列预测基因功能变得愈发可能。

In computational biology, gene prediction or gene finding refers to the process of identifying the regions of genomic DNA that encode genes. This includes protein-coding genes as well as RNA genes, but may also include prediction of other functional elements such as regulatory regions. Gene finding is one of the first and most important steps in understanding the genome of a species once it has been sequenced. In its earliest days, "gene finding" was based on painstaking experimentation on living cells and organisms. Statistical analysis of the rates of homologous recombination of several different genes could determine their order on a certain chromosome, and information from many such experiments could be combined to create a genetic map specifying the rough location of known genes relative to each other. Today, with comprehensive genome sequence and powerful computational resources at the disposal of the research community, gene finding has been redefined as a largely computational problem.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements.

View content license ↗
Microbiology

GenBank

基因序列数据库(GenBank)

基因银行(英语:GenBank,另译基因库、基因数据库)是一个开放获取的序列数据库,对所有公开可利用的核苷酸序列与其翻译的蛋白质进行收集并注释。 此数据库是国际核酸序列数据库协作组织(INSDC)的一部分,由美国国家生物技术信息中心(NCBI)主管,NCBI为美国国立卫生研究院的下属机构。GenBank和它的合作者从全球各个实验室接收了超过百万种生物的数据。 成立三十年来,GenBank数据库成为了最重要的也是最有影响力的生物全领域数据库,其数据正被全球数以百万计的研究人员获取与引用。GenBank中的数据量正以每18个月翻一番的速度持续指数增长,在2013年2月的版本194中,数据库包含有1.62亿个序列,含有1500亿个核苷酸堿基。

The GenBank sequence database is an open access, annotated collection of publicly available nucleotide sequences and their protein translations. GenBank is part of the International Nucleotide Sequence Database Collaboration (INSDC) and is produced and maintained by the National Center for Biotechnology Information (NCBI), a division of the United States National Library of Medicine (NLM), part of the National Institutes of Health (NIH). As of GenBank release 271.0 (April 2026), the database contained 53.90 trillion bases and 6.27 billion sequence records, including 261,460,182 GenBank entries containing 7,289,942,983,522 base pairs of sequence data. The database includes sequences from more than 581,000 formally described species. The database was established in 1982 by Walter Goad and the Los Alamos National Laboratory and has become a central resource for biological research. GenBank is built from direct submissions by individual laboratories as well as bulk submissions from large-scale sequencing projects.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements.

View content license ↗
Microbiology

Gene Ontology

基因本体

基因本体论 (GO) 是一项重要的生物信息学计划,旨在统一所有物种的基因和基因产物属性的表示。更具体地说,该项目旨在:1)维护和开发其基因和基因产品属性的受控词汇表; 2)对基因和基因产物进行注释,并同化和传播注释数据; 3) 提供工具,以便轻松访问项目提供的数据的各个方面,并使用 GO 对实验数据进行功能解释,例如通过富集分析。 GO 是开放生物医学本体这一更大的分类工作的一部分,是 OBO Foundry 的初始候选成员之一。基因命名法侧重于基因和基因产物,而基因本体论侧重于基因和基因产物的功能。

The Gene Ontology (GO) is a major bioinformatics initiative to unify the representation of gene and gene product attributes across all species. More specifically, the project aims to: 1) maintain and develop its controlled vocabulary of gene and gene product attributes; 2) annotate genes and gene products, and assimilate and disseminate annotation data; and 3) provide tools for easy access to all aspects of the data provided by the project, and to enable functional interpretation of experimental data using the GO, for example via enrichment analysis. GO is part of a larger classification effort, the Open Biomedical Ontologies, being one of the Initial Candidate Members of the OBO Foundry. Whereas gene nomenclature focuses on gene and gene products, the Gene Ontology focuses on the function of the genes and gene products.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Gene co-expression network

基因共表达网络

基因共表达网络是一种无向图,每个节点代表基因,如果二者存在明显的共表达关系,就用一个边连接两个节点。 对不同的样本或者不同的实验条件建立基因表达谱后,可以通过查看不同样本间产生相似表达模式的基因对建立基因共表达网络。原因是,两个共表达基因在不同的样本中应以相同模式变化。共同表达的基因是由同一转录控制程序控制、功能相关、同一通路或蛋白结构的组成部分,所以基因共表达网络具有生物学意义。 基因共表达网络不指定共表达关系的方向和类型。然而在基因调控网络中,边是有方向的,代表着反应、变换、互作、激活或者抑制的生化过程。而基因共表达网络并不尝试判定因果关系,边只代表基因之间的相关或者依赖关系。有类似功能或参与统一生物功能的基因会产生很多相互作用,在基因共表达网络中会体现为模块或连接丰富的子图。 基因共表达网络一般是用高通量基因表达谱技术(如微阵列和RNA测序)生成的数据集建立的。

A gene co-expression network (GCN) is an undirected graph, where each node corresponds to a gene, and a pair of nodes is connected with an edge if there is a significant co-expression relationship between them. Having gene expression profiles of a number of genes for several samples or experimental conditions, a gene co-expression network can be constructed by looking for pairs of genes which show a similar expression pattern across samples, since the transcript levels of two co-expressed genes rise and fall together across samples. Gene co-expression networks are of biological interest since co-expressed genes are controlled by the same transcriptional regulatory program, functionally related, or members of the same pathway or protein complex. The direction and type of co-expression relationships are not determined in gene co-expression networks; whereas in a gene regulatory network (GRN) a directed edge connects two genes, representing a biochemical process such as a reaction, transformation, interaction, activation or inhibition.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements.

View content license ↗
Microbiology

Gene set enrichment analysis

基因集富集分析

基因集富集分析 (GSEA)(也称为功能富集分析或通路富集分析)是一种识别在大量基因或蛋白质中过度代表的基因或蛋白质类别的方法,这些基因或蛋白质类别可能与不同的表型(例如不同的生物体生长模式或疾病)相关。该方法使用统计方法来识别显着富集或缺失的基因组。转录组学技术和蛋白质组学结果通常会识别数千个基因,用于分析。研究人员进行高通量实验来产生基因组(例如,在不同条件下差异表达的基因),通常希望检索该基因组的功能图谱,以便更好地了解潜在的生物过程。

Gene set enrichment analysis (GSEA) (also called functional enrichment analysis or pathway enrichment analysis) is a method to identify classes of genes or proteins that are over-represented in a large set of genes or proteins, and may have an association with different phenotypes (e.g. different organism growth patterns or diseases). The method uses statistical approaches to identify significantly enriched or depleted groups of genes. Transcriptomics technologies and proteomics results often identify thousands of genes, which are used for the analysis. Researchers performing high-throughput experiments that yield sets of genes (for example, genes that are differentially expressed under different conditions) often want to retrieve a functional profile of that gene set, in order to better understand the underlying biological processes.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Flux balance analysis

通量平衡分析

在生物化学和系统生物学中,通量平衡分析 (FBA) 是一种利用代谢网络的基因组规模重建来模拟细胞或整个单细胞生物(例如大肠杆菌或酵母)代谢的数学方法。基因组规模重建描述了生物体基于其整个基因组的所有已知或假设的生化反应。这些重建通过关注代谢物之间的相互转化来模拟代谢,识别哪些代谢物参与细胞或生物体中发生的各种反应,并确定编码催化这些反应的酶(如果有)的基因。

In biochemistry and systems biology, flux balance analysis (FBA) is a mathematical method for simulating the metabolism of cells or entire unicellular organisms, such as E. coli or yeast, using genome-scale reconstructions of metabolic networks. Genome-scale reconstructions describe all known or hypothesized biochemical reactions in an organism based on its entire genome. These reconstructions model metabolism by focusing on the interconversions between metabolites, identifying which metabolites are involved in the various reactions taking place in a cell or organism, and determining the genes that encode the enzymes which catalyze these reactions (if any).

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Fluxomics

通量组学

通量组学描述了寻求确定生物实体内代谢反应速率的各种方法。虽然代谢组学可以提供生物样品中代谢物的即时信息,但代谢是一个动态过程。通量组学的意义在于代谢通量决定细胞表型。它的另一个优点是基于代谢组,其成分比基因组或蛋白质组少。 Fluxomics属于随着高通量技术的出现而发展起来的系统生物学领域。系统生物学认识到生物系统的复杂性,并具有解释和预测这种复杂行为的更广泛目标。

Fluxomics describes the various approaches that seek to determine the rates of metabolic reactions within a biological entity. While metabolomics can provide instantaneous information on the metabolites in a biological sample, metabolism is a dynamic process. The significance of fluxomics is that metabolic fluxes determine the cellular phenotype. It has the added advantage of being based on the metabolome which has fewer components than the genome or proteome. Fluxomics falls within the field of systems biology which developed with the appearance of high throughput technologies. Systems biology recognizes the complexity of biological systems and has the broader goal of explaining and predicting this complex behavior.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Genome profiling

基因组分析

基因组分析(GP)是一种无需测序即可获取基因组信息的生物技术。可用于生物体的鉴定和分类。它是由日本生物物理学家Koichi Nishigaki教授及其在埼玉大学的同事于1990年及其后首创的。为了避免混淆,术语“DNA 分析”更改为“基因组分析”,因为术语“DNA 分析”已开始用于法医学领域的不同技术。在 GP 中,基因组 DNA 的小片段被随机扩增(随机 PCR),并对随机 PCR 产物进行温度梯度凝胶电泳 (TGGE),以生成物种特异性迁移模式(基因组图谱)。由此分配物种识别点 (spiddos)。这种方法显然是优越的,因为它不需要任何基因序列的先验知识。

Genome profiling (GP) is a biotechnology that acquires genome information without sequencing. It can be used for identification and classification of organisms. It was pioneered by Japanese biophysicist Prof. Koichi Nishigaki and his colleagues at Saitama University in 1990 and later. The term 'DNA profiling' was changed to 'genome profiling' to avoid confusion, as the term 'DNA profiling' had begun to be used for a different technology in the field of forensics. In GP, small fragments of genomic DNA are randomly amplified (random PCR) and the random PCR products are subjected to temperature-gradient gel electrophoresis (TGGE) to generate a species-specific mobility pattern (genome profile). From this, species identification dots (spiddos) are assigned. This approach is clearly superior because it does not require prior knowledge of any gene sequence.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Genome survey sequence

基因组调查序列

在生物信息学和计算生物学领域,基因组调查序列(GSS)是类似于表达序列标签(EST)的核苷酸序列,唯一的区别是它们大多数起源于基因组,而不是mRNA。基因组调查序列通常由执行基因组测序的实验室生成并提交给 NCBI,并用作标准 GenBank 部门中包含的基因组大小片段的绘图和测序框架等。

In the fields of bioinformatics and computational biology, Genome survey sequences (GSS) are nucleotide sequences similar to expressed sequence tags (ESTs) that the only difference is that most of them are genomic in origin, rather than mRNA. Genome survey sequences are typically generated and submitted to NCBI by labs performing genome sequencing and are used, amongst other things, as a framework for the mapping and sequencing of genome size pieces included in the standard GenBank divisions.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Phylogenetic profiling

系统发育分析

系统发育分析是一种生物信息学技术,其中大量物种中两个性状的联合存在或联合缺失用于推断有意义的生物联系,例如两种不同蛋白质参与同一生物途径。除了检查保守同线性、保守操纵子结构或“罗塞塔石碑”域融合之外,比较系统发育图谱也是一种指定的“后同源”技术,因为该方法所必需的计算在确定哪些蛋白质与哪些蛋白质同源后开始。其中许多技术是由 David Eisenberg 及其同事开发的。系统发育谱比较由 Pellegrini 等人于 1999 年引入。

Phylogenetic profiling is a bioinformatics technique in which the joint presence or joint absence of two traits across large numbers of species is used to infer a meaningful biological connection, such as involvement of two different proteins in the same biological pathway. Along with examination of conserved synteny, conserved operon structure, or "Rosetta Stone" domain fusions, comparing phylogenetic profiles is a designated "post-homology" technique, in that the computation essential to this method begins after it is determined which proteins are homologous to which. A number of these techniques were developed by David Eisenberg and colleagues; phylogenetic profile comparison was introduced in 1999 by Pellegrini, et al.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Neighbor joining

邻接法

在生物信息学中,邻接是一种自下而上(凝聚)的聚类方法,用于创建系统发育树,由 Naruya Saitou 和 Masatoshi Nei 于 1987 年创建。该算法通常基于 DNA 或蛋白质序列数据,需要了解每对类群(例如物种或序列)之间的距离才能创建系统发育树。

In bioinformatics, neighbor joining is a bottom-up (agglomerative) clustering method for the creation of phylogenetic trees, created by Naruya Saitou and Masatoshi Nei in 1987. Usually based on DNA or protein sequence data, the algorithm requires knowledge of the distance between each pair of taxa (e.g., species or sequences) to create the phylogenetic tree.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Bayesian inference in phylogeny

贝叶斯法

系统发育的贝叶斯推理结合了先验和数据似然中的信息来创建所谓的树的后验概率,即给定数据、先验和似然模型时树正确的概率。贝叶斯推理在 20 世纪 90 年代被三个独立的小组引入分子系统发育学:伯克利的 Bruce Rannala 和 Ziheng Yang、麦迪逊的 Bob Mau 以及爱荷华大学的 Shuying Li,最后两位当时是博士生。自 2001 年 MrBayes 软件发布以来,该方法变得非常流行,现在是分子系统发育学中最流行的方法之一。

Bayesian inference of phylogeny combines the information in the prior and in the data likelihood to create the so-called posterior probability of trees, which is the probability that the tree is correct given the data, the prior and the likelihood model. Bayesian inference was introduced into molecular phylogenetics in the 1990s by three independent groups: Bruce Rannala and Ziheng Yang in Berkeley, Bob Mau in Madison, and Shuying Li in University of Iowa, the last two being PhD students at the time. The approach has become very popular since the release of the MrBayes software in 2001, and is now one of the most popular methods in molecular phylogenetics.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Hidden Markov model

隐马尔可夫模型

在概率论中,隐马尔可夫模型(HMM)是一种马尔可夫模型,其中观测值依赖于潜在(或隐藏)马尔可夫过程(称为 X {\displaystyle X} )。 HMM 要求存在一个可观察过程 Y {\displaystyle Y},其结果以已知方式取决于 X {\displaystyle X} 的结果。由于 X {\displaystyle X} 无法直接观察,因此目标是通过观察 Y {\displaystyle Y} 来了解 X {\displaystyle X} 的状态。

In probability theory, a hidden Markov model (HMM) is a Markov model in which the observations are dependent on a latent (or hidden) Markov process (referred to as X {\displaystyle X} ). An HMM requires that there be an observable process Y {\displaystyle Y} whose outcomes depend on the outcomes of X {\displaystyle X} in a known way. Since X {\displaystyle X} cannot be observed directly, the goal is to learn about state of X {\displaystyle X} by observing Y {\displaystyle Y} .

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Heat map

熱圖

热图(或热图)是一种二维数据可视化技术,它将数据集中各个值的大小表示为颜色。颜色的变化可以是色调或强度。在某些应用程序(例如犯罪分析或网站点击跟踪)中,颜色用于表示数据点的密度,而不是与每个点相关的值。 “热图”是一个相对较新的术语,但着色矩阵的实践已经存在了一个多世纪。

A heat map (or heatmap) is a two-dimensional data visualization technique that represents the magnitude of individual values within a dataset as a color. The variation in color may be by hue or intensity. In some applications such as crime analytics or website click-tracking, color is used to represent the density of data points rather than a value associated with each point. "Heat map" is a relatively new term, but the practice of shading matrices has existed for over a century.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

False discovery rate

偽發現率

假发现率(False discovery rate, FDR)完善了对多重假设测试的检验, F D R = Q e = E [ Q ] , {\displaystyle \mathrm {FDR} =Q_{e}=\mathrm {E} \!\left[Q\right],} 其中E表示期望, Q = V / R = V / ( V + S ) {\displaystyle Q=V/R=V/(V+S)} ,V表示错误拒绝零假设的数目,R表示拒绝零假设的数目。R取0时FDR直接取0,写成一句话就是 F D R = E [ V / R | R > 0 ] ⋅ P (…

In statistics, the false discovery rate (FDR) is a method of conceptualizing the rate of type I errors in null hypothesis testing when conducting multiple comparisons. FDR-controlling procedures are designed to control the FDR, which is the expected proportion of "discoveries" (rejected null hypotheses) that are false (incorrect rejections of the null). Equivalently, the FDR is the expected ratio of the number of false positive classifications (false discoveries) to the total number of positive classifications (rejections of the null). The total number of rejections of the null include both the number of false positives (FP) and true positives (TP). Simply put, FDR = FP / (FP + TP). FDR-controlling procedures provide less stringent control of Type I errors compared to family-wise error rate (FWER) controlling procedures (such as the Bonferroni correction), which control the probability of at least one Type I error. Thus, FDR-controlling procedures have greater power, at the cost of increased numbers of Type I errors.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements.

View content license ↗
Microbiology

Synthetic biology

合成生物学

合成生物学(英语:synthetic biology)是将生物科学应用到日常生活中的一种崭新方式。英国伦敦的皇家科学院(Royal Society)认为:合成生物学结合了其他领域的知识与工具,涉及的领域包括系统生物学、基因工程、机械工程、机电工程、信息论、物理学、纳米技术及电脑模拟等等。 目前,合成生物学已在多个行业落实应用,例如农业、能源、制造业及医学等等。

Synthetic biology (SynBio) is a multidisciplinary scientific field that applies the principles of engineering to develop new biological parts, devices, and systems or to redesign existing systems found in nature. The field encompasses a broad range of methodologies from various disciplines, such as biochemistry, biotechnology, biomaterials, material science/engineering, genetic engineering, molecular biology, molecular engineering, systems biology, membrane science, biophysics, chemical and biological engineering, electrical and computer engineering, control engineering and evolutionary biology. It includes designing and constructing biological modules, biological systems, and biological machines, or re-designing existing biological systems for useful purposes. Additionally, it is the branch of science that focuses on the new abilities of engineering into existing organisms to redesign them for useful purposes.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements.

View content license ↗
Microbiology

Genomics

基因組學

基因组学是分子生物学的一个跨学科领域,专注于基因组的结构、功能、进化、作图和编辑。基因组是生物体的完整 DNA 集,包括其所有基因及其分层的三维结构配置。遗传学是指研究个体基因及其在遗传中的作用,与此相反,基因组学旨在对生物体的所有基因、它们的相互关系以及对生物体的影响进行集体表征和量化。基因可以在酶和信使分子的帮助下指导蛋白质的产生。反过来,蛋白质构成器官和组织等身体结构,并控制化学反应并在细胞之间传递信号。

Genomics is an interdisciplinary field of molecular biology focusing on the structure, function, evolution, mapping, and editing of genomes. A genome is an organism's complete set of DNA, including all of its genes as well as its hierarchical, three-dimensional structural configuration. In contrast to genetics, which refers to the study of individual genes and their roles in inheritance, genomics aims at the collective characterization and quantification of all of an organism's genes, their interrelations and influence on the organism. Genes may direct the production of proteins with the assistance of enzymes and messenger molecules. In turn, proteins make up body structures such as organs and tissues as well as control chemical reactions and carry signals between cells.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Transcriptomics technologies

转录组学技术

转录组学技术是用于研究生物体转录组(其所有 RNA 转录本的总和)的技术。生物体的信息内容记录在其基因组的DNA中并通过转录表达。在这里,mRNA 充当信息网络中的瞬时中间分子,而非编码 RNA 则执行其他不同的功能。转录组捕获细胞中存在的总转录本的及时快照。转录组学技术广泛描述了哪些细胞过程是活跃的,哪些是休眠的。分子生物学的一个主要挑战是了解单个基因组如何产生多种细胞。另一个是基因表达的调控方式。研究整个转录组的第一次尝试始于 20 世纪 90 年代初。

Transcriptomics technologies are the techniques used to study an organism's transcriptome, the sum of all of its RNA transcripts. The information content of an organism is recorded in the DNA of its genome and expressed through transcription. Here, mRNA serves as a transient intermediary molecule in the information network, whilst non-coding RNAs perform additional diverse functions. A transcriptome captures a snapshot in time of the total transcripts present in a cell. Transcriptomics technologies provide a broad account of which cellular processes are active and which are dormant. A major challenge in molecular biology is to understand how a single genome gives rise to a variety of cells. Another is how gene expression is regulated. The first attempts to study whole transcriptomes began in the early 1990s.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗
Microbiology

Systems biology

系统生物学

系统生物学是复杂生物系统的计算和数学分析和建模。它是一个基于生物学的跨学科研究领域,专注于生物系统内复杂的相互作用,使用整体方法(整体论而不是更传统的还原论)进行生物学研究。这个多方面的研究领域需要化学家、生物学家、数学家、物理学家和工程师的共同努力,通过将各种定量分子测量与精心构建的数学模型相结合来破译复杂生命系统的生物学。它代表了理解生物系统内复杂关系的综合方法。

Systems biology is the computational and mathematical analysis and modeling of complex biological systems. It is a biology-based interdisciplinary field of study that focuses on complex interactions within biological systems, using a holistic approach (holism instead of the more traditional reductionism) to biological research. This multifaceted research domain necessitates the collaborative efforts of chemists, biologists, mathematicians, physicists, and engineers to decipher the biology of intricate living systems by merging various quantitative molecular measurements with carefully constructed mathematical models. It represents a comprehensive method for comprehending the complex relationships within biological systems.

Sources, licensing and use

Wikipedia contributors · Retrieved2026-10-04 · CC BY-SA 4.0. Introductions were extracted as plain text and shortened. Language versions may emphasize different aspects.For concept reference; consult the original standards for authoritative requirements. The Chinese definition is a machine-assisted translation of the cited English introduction; check technical terminology against the original.

View content license ↗

How can this knowledge be incorporated into high-end products?

Relevant scientific figures and methodological contributions

Understand these concepts in the tool

This collection combines attributed Wikipedia excerpts and original SciAtlas bilingual definitions under CC BY-SA 4.0. Excerpts were extracted and shortened; machine-assisted Chinese translations are labeled. Original entries provide further reading. Language versions may differ in emphasis and do not replace standards. Concepts can appear in several disciplines; consult standards and original literature for rigorous use.

Knowledge Snapshot: 2026-10-04. Category cross-inclusion is used for reading navigation, and the names of people, institutions and unexplained placeholders in the field are not included in the quantity.