![]() |
OpenMS
|
Representation of an experimental design in OpenMS. Instances can be loaded with the ExperimentalDesignFile class. More...
#include <OpenMS/METADATA/ExperimentalDesign.h>
Classes | |
| class | MSFileSectionEntry |
| One row of the MS file section: one quantitative channel of one MS file. More... | |
| class | SampleSection |
| The sample section: one named row per sample, one column per factor. More... | |
Public Types | |
| using | MSFileSection = std::vector< MSFileSectionEntry > |
Public Member Functions | |
| ExperimentalDesign ()=default | |
| ExperimentalDesign (const MSFileSection &msfile_section, const SampleSection &sample_section) | |
| const MSFileSection & | getMSFileSection () const |
| void | setMSFileSection (const MSFileSection &msfile_section) |
| const ExperimentalDesign::SampleSection & | getSampleSection () const |
| void | setSampleSection (const SampleSection &sample_section) |
| std::map< std::vector< std::string >, std::set< std::string > > | getUniqueSampleRowToSampleMapping () const |
| std::map< std::string, unsigned > | getSampleToPrefractionationMapping () const |
| std::map< unsigned int, std::vector< std::string > > | getFractionToMSFilesMapping () const |
| return fraction index to file paths (ordered by fraction_group) | |
| std::vector< std::vector< std::pair< std::string, unsigned > > > | getConditionToPathLabelVector () const |
| std::map< std::vector< std::string >, std::set< unsigned > > | getConditionToSampleMapping () const |
| return a condition to Sample index mapping | |
| std::map< std::pair< std::string, unsigned >, unsigned > | getPathLabelToPrefractionationMapping (bool use_basename_only) const |
| std::map< std::pair< std::string, unsigned >, unsigned > | getPathLabelToConditionMapping (bool use_basename_only) const |
| std::map< std::string, unsigned > | getSampleToConditionMapping () const |
| std::map< std::pair< std::string, unsigned >, unsigned > | getPathLabelToSampleMapping (bool use_basename_only) const |
| return <file_path, label> to sample index mapping | |
| std::map< std::pair< std::string, unsigned >, unsigned > | getPathLabelToFractionMapping (bool use_basename_only) const |
| return <file_path, label> to fraction mapping | |
| std::map< std::pair< std::string, unsigned >, unsigned > | getPathLabelToFractionGroupMapping (bool use_basename_only) const |
| return <file_path, label> to fraction_group mapping | |
| unsigned | getNumberOfSamples () const |
| Number of samples measured (= number of rows in the sample section) | |
| unsigned | getNumberOfFractions () const |
| Number of distinct fraction indices used anywhere in the design. | |
| unsigned | getNumberOfLabels () const |
| Highest label index used anywhere in the design. | |
| unsigned | getNumberOfMSFiles () const |
| Number of distinct MS file paths in the design. | |
| unsigned | getNumberOfFractionGroups () const |
| Number of distinct fraction groups. | |
| unsigned | getSample (unsigned fraction_group, unsigned label=1) |
| Sample quantified in a given fraction group and label. | |
| bool | isFractionated () const |
| Whether the design is fractionated. | |
| Size | filterByBasenames (const std::set< std::string > &bns) |
| bool | sameNrOfMSFilesPerFraction () const |
| Size | annotateColumnHeaders (ConsensusMap &cmap) const |
| Write this design's fraction structure onto a ConsensusMap's column headers. | |
Static Public Member Functions | |
| static ExperimentalDesign | fromConsensusMap (const ConsensusMap &c) |
| Extract experimental design from consensus map. | |
| static ExperimentalDesign | fromFeatureMap (const FeatureMap &f) |
| Extract experimental design from feature map. | |
| static ExperimentalDesign | fromIdentifications (const std::vector< ProteinIdentification > &proteins) |
| Extract experimental design from identifications. | |
Private Member Functions | |
| std::map< Size, std::string > | sampleRowToName_ () const |
| std::vector< std::string > | getFileNames_ (bool basename) const |
| std::vector< unsigned > | getLabels_ () const |
| std::vector< unsigned > | getFractions_ () const |
| std::map< std::pair< std::string, unsigned >, unsigned > | pathLabelMapper_ (bool, unsigned(*f)(const ExperimentalDesign::MSFileSectionEntry &)) const |
| Generic Mapper (Path, Label) -> f(row) | |
| void | sort_ () |
| void | isValid_ () |
Static Private Member Functions | |
| template<typename T > | |
| static void | errorIfAlreadyExists (std::set< T > &container, T &item, const std::string &message) |
Private Attributes | |
| MSFileSection | msfile_section_ |
| SampleSection | sample_section_ |
Representation of an experimental design in OpenMS. Instances can be loaded with the ExperimentalDesignFile class.
An experimental design maps quantitative values to the biological material they were measured from, and records the metadata needed to compare those measurements. ProteomicsLFQ, IsobaricWorkflow, ProteinQuantifier and MSstatsConverter all take one.
It is a TAB-separated text file (conventionally design.tsv) with a few numeric columns describing the measurement layout plus any number of free-form columns describing the biology.
One row describes one measured quantity: one channel of one MS file. A label-free run has one channel and contributes one row; a TMT10-plex run has ten and contributes ten rows – the same file name ten times, with Label 1 to 10.
Per row:
| Column | Meaning |
|---|---|
Spectra_Filepath | The file that was measured |
Label | The channel inside that file |
Fraction | Which slice of the injected material the file contains |
Fraction_Group | Which rows are combined into one complete measurement |
Sample | The biological material measured in this row |
These five columns make up the MS file section. The first four are validated strictly (see Rules OpenMS enforces). Sample links to the sample section: a free-form table of columns – OpenMS calls them factors – such as condition, biological replicate, genotype or time point. OpenMS does not interpret factor values; it only compares them to decide which samples belong together.
Lab usage of these terms varies, so each entry gives the OpenMS meaning, how it is written in the file, and what OpenMS does with it.
MS file (run) – one raw/mzML file, i.e. one LC-MS measurement. Named by Spectra_Filepath. A file appears in as many rows as it has channels.
Sample – the biological material a quantity is measured from; concretely, a named row of the sample section that Sample in the file section points at. Two file rows with the same Sample value describe measurements of the same material; rows with different values are different material, even if all their metadata columns agree. Same-sample rows within one fraction group are fractions of one measurement and are combined into one quantity; same-sample rows in different fraction groups stay separate quantities, never pooled by OpenMS. In a labeled experiment each channel normally names a different sample, so one TMT10 file usually describes ten, but repeating a Sample across channels or mixtures is legal and is how a bridge/reference channel is declared. getNumberOfSamples() returns the number of rows in the sample section.
Sample holds a name, i.e. an arbitrary string – sdrf-pipelines writes "1", "2", ..., but "BSA1" or "patient_7" are equally valid. In the C++ API, MSFileSectionEntry::sample_name is that name while MSFileSectionEntry::sample is the zero-based row index of the sample in the sample section. The two rarely coincide; do not use one where the other is expected.Fraction – one portion of a sample that was separated before LC-MS (high-pH reversed phase, SCX, gel bands, ...) and measured in its own file. Fractions are parts of one measurement rather than repeats of it: tools that report one value per assay (ProteinQuantifier, mzTab, QPX) combine them – summed by default, or reduced to one fraction with fractions:aggregate best. The MSstats export does not: it emits one row per fraction and lets MSstats summarize. Numbered from 1; use 1 in every row for unfractionated data. Use the same numbers in every fraction group so corresponding fractions line up (getFractionToMSFilesMapping() groups by this number across groups).
Fraction group – the rows that form one complete measurement of one injected mixture: all fractions combined before a quantity is reported. The unit a value is actually reported for is the pair (Fraction_Group, Label) – ProteinQuantifier and mzTab call it an assay, the QPX exporters a quantification unit. A label-free fraction group therefore yields one quantity, a TMT10 group ten. Must be integers starting at 1, consecutive, no gaps. Unfractionated label-free data has one fraction group per file. Measuring the same material again means a new fraction group with the same Sample; see technical replicate below.
Label (channel) – the multiplexing channel within one file. 1 for label-free and DIA data; 1..n for an n-plex, 1-based in the canonical channel order of the reagent (TMT10-plex: 126→1, 127N→2, 127C→3, 128N→4, 128C→5, 129N→6, 129C→7, 130N→8, 130C→9, 131→10; SILAC 2-plex: light→1, heavy→2; iTRAQ4-plex: 114→1 ... 117→4). The value is a channel position, not a reagent name: OpenMS resolves it to a reagent through the quantification method that produced the data, so a mismatch swaps channels without any error. Every (file, label) pair may occur only once in a design.
Technical replicate – the same material measured again (re-injection, re-run of the same digest). Written as a second fraction group with the same Sample value. OpenMS keeps the repeats as separate quantities; collapsing them is left to the statistics package.
Biological replicate – an independent biological unit (another animal, patient, culture) measured under a condition. Different biological replicates are different samples; the relationship is declared in the sample section, conventionally in a column called MSstats_BioReplicate.
Conditions and factors belong to the sample section, see The sample section: factors and conditions.
Both are TAB-separated (tabs only – aligning columns with spaces does not work). While parsing, cells are whitespace trimmed and lines starting with a hash character are ignored.
One-table format – a single table. Mandatory columns are Fraction_Group, Fraction and Spectra_Filepath; Label and Sample are optional; any further column is taken to be sample metadata. This is the more common form.
Two-table format – an MS file section and a sample section, separated by at least one blank line. The file section accepts only Fraction_Group, Fraction, Spectra_Filepath, Label and Sample; any other column there is a parse error. The sample section must have a Sample column and may carry any number of further columns. Useful when many files share a sample, to avoid repeating its metadata.
The format is auto-detected, and the detector is cruder than the parsers: it scans every line, not just headers, and reads the file as two-table as soon as one line has exactly one cell equal to Sample and no cell equal to Fraction_Group. It does not trim cells or skip comment lines, so a sample header whose cells carry padding is missed and a commented-out line still counts.
Six independent biological units, one file each. Every file is its own fraction group and its own sample; the fraction column is 1 throughout because nothing was fractionated.
| Fraction_Group | Fraction | Spectra_Filepath | Label | Sample | MSstats_Condition | MSstats_BioReplicate |
|---|---|---|---|---|---|---|
| 1 | 1 | ctrl_1.mzML | 1 | 1 | control | 1 |
| 2 | 1 | ctrl_2.mzML | 1 | 2 | control | 2 |
| 3 | 1 | ctrl_3.mzML | 1 | 3 | control | 3 |
| 4 | 1 | drug_1.mzML | 1 | 4 | treated | 4 |
| 5 | 1 | drug_2.mzML | 1 | 5 | treated | 5 |
| 6 | 1 | drug_3.mzML | 1 | 6 | treated | 6 |
6 files, 6 fraction groups, 6 samples, 1 fraction, 1 label, 2 conditions.
Two patients, each digest injected twice. The repeats carry the same Sample and metadata, but their own fraction group:
| Fraction_Group | Fraction | Spectra_Filepath | Label | Sample | MSstats_Condition | MSstats_BioReplicate |
|---|---|---|---|---|---|---|
| 1 | 1 | p1_inj1.mzML | 1 | 1 | control | 1 |
| 2 | 1 | p1_inj2.mzML | 1 | 1 | control | 1 |
| 3 | 1 | p2_inj1.mzML | 1 | 2 | treated | 2 |
| 4 | 1 | p2_inj2.mzML | 1 | 2 | treated | 2 |
4 files, 4 fraction groups, 2 samples. Reusing a sample across fraction groups is allowed and does not merge them: OpenMS reports four quantities and leaves it to the statistics package to treat them as repeated measurements.
Two samples again, as in example 2, but six files instead of four:
| Fraction_Group | Fraction | Spectra_Filepath | Label | Sample | MSstats_Condition | MSstats_BioReplicate |
|---|---|---|---|---|---|---|
| 1 | 1 | ctrl_F1.mzML | 1 | 1 | control | 1 |
| 1 | 2 | ctrl_F2.mzML | 1 | 1 | control | 1 |
| 1 | 3 | ctrl_F3.mzML | 1 | 1 | control | 1 |
| 2 | 1 | drug_F1.mzML | 1 | 2 | treated | 2 |
| 2 | 2 | drug_F2.mzML | 1 | 2 | treated | 2 |
| 2 | 3 | drug_F3.mzML | 1 | 2 | treated | 2 |
6 files, 2 fraction groups, 2 samples, 3 fractions. In example 2 four files produced four quantities; here six files produce two, because the files of a fraction group are aggregated. Fraction numbers repeat across groups deliberately: ctrl_F2 and drug_F2 are the same slice of the separation and are aligned with each other.
One file, ten channels, ten samples. The file name repeats on every row; of the measurement columns only Label changes, and each channel names its own sample.
| Fraction_Group | Fraction | Spectra_Filepath | Label | Sample | MSstats_Condition | MSstats_BioReplicate | MSstats_Mixture |
|---|---|---|---|---|---|---|---|
| 1 | 1 | tmt_mix1.mzML | 1 | 1 | control | 1 | 1 |
| 1 | 1 | tmt_mix1.mzML | 2 | 2 | control | 2 | 1 |
| 1 | 1 | tmt_mix1.mzML | 3 | 3 | control | 3 | 1 |
| 1 | 1 | tmt_mix1.mzML | 4 | 4 | control | 4 | 1 |
| 1 | 1 | tmt_mix1.mzML | 5 | 5 | control | 5 | 1 |
| 1 | 1 | tmt_mix1.mzML | 6 | 6 | treated | 6 | 1 |
| 1 | 1 | tmt_mix1.mzML | 7 | 7 | treated | 7 | 1 |
| 1 | 1 | tmt_mix1.mzML | 8 | 8 | treated | 8 | 1 |
| 1 | 1 | tmt_mix1.mzML | 9 | 9 | treated | 9 | 1 |
| 1 | 1 | tmt_mix1.mzML | 10 | 10 | treated | 10 | 1 |
1 file, 1 fraction group, 10 labels, 10 samples. Label 1 is TMT channel 126 and Label 10 is channel 131 – see the glossary for the full order.
Two TMT mixtures, each separated into three fractions: 2 × 3 = 6 files, 6 × 10 = 60 rows, of which a few representative ones are shown.
| Fraction_Group | Fraction | Spectra_Filepath | Label | Sample | MSstats_Condition | MSstats_BioReplicate | MSstats_Mixture |
|---|---|---|---|---|---|---|---|
| 1 | 1 | mix1_F1.mzML | 1 | 1 | control | 1 | 1 |
| ... | ... | ... | ... | ... | ... | ... | ... |
| 1 | 1 | mix1_F1.mzML | 10 | 10 | treated | 10 | 1 |
| 1 | 2 | mix1_F2.mzML | 1 | 1 | control | 1 | 1 |
| ... | ... | ... | ... | ... | ... | ... | ... |
| 1 | 3 | mix1_F3.mzML | 10 | 10 | treated | 10 | 1 |
| 2 | 1 | mix2_F1.mzML | 1 | 11 | control | 11 | 2 |
| ... | ... | ... | ... | ... | ... | ... | ... |
| 2 | 3 | mix2_F3.mzML | 10 | 20 | treated | 20 | 2 |
6 files, 2 fraction groups, 3 fractions, 10 labels, 20 samples. A sample repeats across the fractions of its own group (same material, different slice) but not across the two mixtures (different material, different TMT tube).
MSstats_Mixture is an ordinary factor – its name contains no "replicate" – so it counts towards the condition. This design therefore has four OpenMS conditions (control/treated x mixture 1/2), not two, and the mixtures never share one. That affects only OpenMS' own condition grouping; the column is forwarded to MSstats unchanged.Metabolic labeling: two channels per file, two files.
| Fraction_Group | Fraction | Spectra_Filepath | Label | Sample | MSstats_Condition | MSstats_BioReplicate |
|---|---|---|---|---|---|---|
| 1 | 1 | silac_r1.mzML | 1 | 1 | control | 1 |
| 1 | 1 | silac_r1.mzML | 2 | 2 | treated | 2 |
| 2 | 1 | silac_r2.mzML | 1 | 3 | control | 3 |
| 2 | 1 | silac_r2.mzML | 2 | 4 | treated | 4 |
Label 1 is the light channel, Label 2 the heavy one. In a 3-plex, medium is 2 and heavy is 3.
-method LFQ refuses a design with more than one label and -method ISO requires MSstats_Mixture, and ProteomicsLFQ refuses multi-label designs outright. ProteinQuantifier consumes them. The MSstats columns above are shown only for consistency with the other examples.Example 3 with the metadata factored out. The blank line separating the sections is mandatory.
MS file section:
| Fraction_Group | Fraction | Spectra_Filepath | Label | Sample |
|---|---|---|---|---|
| 1 | 1 | ctrl_F1.mzML | 1 | 1 |
| 1 | 2 | ctrl_F2.mzML | 1 | 1 |
| 1 | 3 | ctrl_F3.mzML | 1 | 1 |
| 2 | 1 | drug_F1.mzML | 1 | 2 |
| 2 | 2 | drug_F2.mzML | 1 | 2 |
| 2 | 3 | drug_F3.mzML | 1 | 2 |
(blank line)
Sample section:
| Sample | MSstats_Condition | MSstats_BioReplicate |
|---|---|---|
| 1 | control | 1 |
| 2 | treated | 2 |
Every column of the sample section is a factor. Names are free-form; OpenMS only compares values, so Genotype, Timepoint or Dose behave like the MSstats columns below.
A condition is a unique combination of the values of all factors except Sample and except every factor whose column name contains "replicate" or "Replicate". Samples that agree on all remaining factors form one condition. For a design whose sample section has factor columns, this is what getConditionToSampleMapping(), getSampleToConditionMapping(), getPathLabelToConditionMapping() and getConditionToPathLabelVector() return. For a factor-less section – as built by fromConsensusMap() and fromIdentifications() – they currently disagree: getSampleToConditionMapping() gives every sample its own condition, while the other three collapse all samples into one. Epifany uses getConditionToPathLabelVector() to merge ID runs per condition when a design is supplied; ProteomicsLFQ does not – it merges every ID run study-wide regardless of condition.
Condition numbers are the lexicographic rank of the factor-value tuple, so condition 0 is whichever value sorts first, not a reference level.
MSstats_BioReplicate and Technical_Replicate are excluded from the condition automatically; Donor or Rep are not, and every distinct value in them creates its own condition. This catches MSstats_Mixture too (see example 5). Name replicate columns accordingly.getUniqueSampleRowToSampleMapping() and getSampleToPrefractionationMapping() apply the weaker rule: they ignore only Sample and keep the replicate columns, so they group samples whose metadata rows are completely identical.
The conventional MSstats columns, understood by MSstatsConverter, are:
MSstats_Condition: the condition of a sample (e.g. control, "1000 mMol"). Forwarded to MSstats, where it is used to formulate test contrasts.MSstats_BioReplicate: identifier of the biological unit. A value that recurs under two different conditions declares one unit measured in both, i.e. a paired design; give unrelated samples distinct values.MSstats_Mixture (isobaric labeling only): identifier of the labeled mixture that was measured together in one tube. Samples labeled with different TMT reagents of the same mix share a mixture identifier; technical replicates of a mixture keep it as well.See the MSstats manual for what MSstats does with them.
Checked when the file is loaded; a design that breaks one of these is rejected:
(Fraction_Group, Fraction, Label) triple may occur only once.(Spectra_Filepath, Label) pair may occur only once.Label value (normally label-free), each (Fraction_Group, Label) pair must map to exactly one sample. Two samples in one such fraction group means the labels are missing or the grouping is wrong.Sample becomes mandatory as soon as any Label is greater than 1 – OpenMS cannot guess which channel is which material. Without a Sample column, the sample name defaults to the Fraction_Group value, which makes every fraction group its own sample.Sample value used in the file section must exist in the sample section, otherwise loading fails with a bare std::out_of_range. Omitting the Sample column from a two-table file section therefore fails as well, unless the sample section literally contains rows named "Fraction group 1", "Fraction group 2", ...Spectra_Filepath is resolved first against the directory of the design file, then against the current working directory; if neither exists the string is kept as written. Tools that need the spectra present (require_spectra_files) fail instead.Sample is the dangerous case: sample names then fall back to the fraction-group value, so reused samples quietly become distinct ones and the misspelled column also splits conditions. The two-table file section rejects unknown headers.SDRF-Proteomics is the HUPO-PSI sample-and-data-relationship format used by PRIDE and quantms. It covers more than an OpenMS design (organism, disease, instrument, search settings) and can generate one: parse_sdrf convert-openms -s sdrf.tsv from sdrf-pipelines writes the file described here. The terms correspond as follows – for the exact conversion, which varies between releases, consult sdrf-pipelines itself:
| OpenMS design | SDRF-Proteomics | Notes |
|---|---|---|
Spectra_Filepath | comment[data file] | copied verbatim; the converter can optionally rewrite the extension, e.g. .raw to .mzML |
Fraction | comment[fraction identifier] | 1 when absent or "not available" |
Label | comment[label] | the ontology term becomes its 1-based index within the plex inferred for that file, so the same reagent name can map to different numbers in different plexes |
Fraction_Group | source name + comment[technical replicate] | assigned per MS file, then renumbered to stay consecutive. Label-free data gets one group per (source, technical replicate) – which is why re-injections get their own group; in a labeled design all channels of a file necessarily share its group |
Sample | source name | a source name ending in "sample N" contributes N, otherwise names are numbered in order of appearance |
MSstats_Condition | the factor value[...] columns, joined with a pipe | falls back to the non-redundant characteristics[...] columns, then to the source name |
MSstats_BioReplicate | derived from source name | the source's sample number, so equal to Sample in a uniform design; not copied from characteristics[biological replicate] |
MSstats_Mixture | derived per labeled file and sample | isobaric designs only (TMT/iTRAQ); SILAC gets no mixture column |
Two SDRF terms do not map one-to-one. comment[technical replicate] is folded into Fraction_Group, since OpenMS expresses technical replication as "same sample, new fraction
group" rather than as a column. And SDRF's source name (the material) vs. assay name (the measurement) corresponds conceptually to Sample vs. the file/label pair, though the OpenMS converter reads only the former.
Label 1 for every SILAC channel, which OpenMS then rejects as a duplicate (Spectra_Filepath, Label) pair.QPX ("Quantitative Proteomics eXchange") is the Parquet exchange format written via the out_qpx option of ProteomicsLFQ, IsobaricWorkflow and ProteinQuantifier. OpenMS writes three views – quantms.feature.parquet, quantms.psm.parquet and quantms.pg.parquet. It does not write run.parquet or sample.parquet; those are generated from the SDRF by the quantms converter.
What the design controls in the views OpenMS writes:
| OpenMS design | QPX | Notes |
|---|---|---|
basename of Spectra_Filepath, no extension | run_file_name | primary-key component of the psm and feature views; the pg view instead carries the run names in grouped_runs. See ArrowIOHelpers::qpxRunFileName() |
Fraction_Group | grouped_runs of the pg view | QPX has no fraction-group number: the group is represented by listing its run names |
Label | intensities[].label (feature) and the scalar label (pg) | as a canonical channel token ("LFQ", "TMT126", "SILAC heavy"), not as the number; see ArrowIOHelpers::qpxIntensityLabels() |
Those labels and run names are join keys against run.parquet, so the design also has to agree with the views OpenMS does not write: Fraction with run.fraction, Sample with run.samples[].sample_accession (keyed by the SDRF source name), and MSstats_BioReplicate with run.samples[].biological_replicate. QPX's run.samples[].technical_replicate has no counterpart in an OpenMS design, because technical replication is expressed as a second fraction group.
(any run of the group, label) resolves to one sample. Ragged designs, a label meaning two samples within one group, and designs that put one run into several fraction groups are refused.(file, channel) has no design row.| using MSFileSection = std::vector<MSFileSectionEntry> |
|
default |
| ExperimentalDesign | ( | const MSFileSection & | msfile_section, |
| const SampleSection & | sample_section | ||
| ) |
| Size annotateColumnHeaders | ( | ConsensusMap & | cmap | ) | const |
Write this design's fraction structure onto a ConsensusMap's column headers.
The inverse of fromConsensusMap(): stamps fraction_group, fraction and sample_name as meta values on every column header whose (basename, label) pair matches a design row. Headers are matched on the 1-based label, so this works for both label-free maps (one header per file) and multiplexed ones (one header per file and channel).
Without this the fraction structure lives only in the design object, and consumers that read the map – exporters, converters – cannot tell two fractions of one sample from two independent runs.
| [in,out] | cmap | Map whose column headers are annotated in place |
|
staticprivate |
| Size filterByBasenames | ( | const std::set< std::string > & | bns | ) |
filters the MSFileSection to only include a given subset of files whose basenames are given with bns
|
static |
Extract experimental design from consensus map.
|
static |
Extract experimental design from feature map.
|
static |
Extract experimental design from identifications.
| std::vector< std::vector< std::pair< std::string, unsigned > > > getConditionToPathLabelVector | ( | ) | const |
return vector of filepath/label combinations that share the same conditions after removing replicate columns in the sample section (e.g. for merging across replicates)
| std::map< std::vector< std::string >, std::set< unsigned > > getConditionToSampleMapping | ( | ) | const |
return a condition to Sample index mapping
A condition is the unique combination of the sample section's factor values, ignoring the Sample column and every column whose NAME contains "replicate" or "Replicate". That naming convention is the only thing that marks a column as a replicate column: a column called MSstats_BioReplicate is excluded, one called Donor is not and would split the replicates it names into separate conditions.
|
private |
|
private |
| std::map< unsigned int, std::vector< std::string > > getFractionToMSFilesMapping | ( | ) | const |
return fraction index to file paths (ordered by fraction_group)
|
private |
| const MSFileSection & getMSFileSection | ( | ) | const |
Referenced by QuantificationUnits::QuantificationUnits(), and MS1LabeledRatioQuantifier::run().
| unsigned getNumberOfFractionGroups | ( | ) | const |
Number of distinct fraction groups.
Allows to group fraction ids and source files.
| unsigned getNumberOfFractions | ( | ) | const |
Number of distinct fraction indices used anywhere in the design.
Meaningful only when all fraction groups use the same fraction numbering, which this does not check. sameNrOfMSFilesPerFraction() is a necessary but not sufficient sanity check: it only verifies that every fraction index occurs in equally many rows.
| unsigned getNumberOfLabels | ( | ) | const |
Highest label index used anywhere in the design.
The plex size for a well-formed design, 1 for label-free. Contiguous labels are not enforced.
| unsigned getNumberOfMSFiles | ( | ) | const |
Number of distinct MS file paths in the design.
A multiplexed file contributes one file but several rows.
| unsigned getNumberOfSamples | ( | ) | const |
Number of samples measured (= number of rows in the sample section)
NOT the highest sample index: MSFileSectionEntry::sample is zero-based, so the highest index is one less than this.
Referenced by QuantificationUnits::QuantificationUnits().
| std::map< std::pair< std::string, unsigned >, unsigned > getPathLabelToConditionMapping | ( | bool | use_basename_only | ) | const |
return <file_path, label> to condition mapping (a condition is a unique combination of all columns in the sample section, except for replicates.
| std::map< std::pair< std::string, unsigned >, unsigned > getPathLabelToFractionGroupMapping | ( | bool | use_basename_only | ) | const |
return <file_path, label> to fraction_group mapping
| std::map< std::pair< std::string, unsigned >, unsigned > getPathLabelToFractionMapping | ( | bool | use_basename_only | ) | const |
return <file_path, label> to fraction mapping
| std::map< std::pair< std::string, unsigned >, unsigned > getPathLabelToPrefractionationMapping | ( | bool | use_basename_only | ) | const |
return <file_path, label> to prefractionation mapping (a prefractionation group is a unique combination of all columns in the sample section, except for replicates.
| std::map< std::pair< std::string, unsigned >, unsigned > getPathLabelToSampleMapping | ( | bool | use_basename_only | ) | const |
return <file_path, label> to sample index mapping
| unsigned getSample | ( | unsigned | fraction_group, |
| unsigned | label = 1 |
||
| ) |
Sample quantified in a given fraction group and label.
fraction_group at label | Exception::ElementNotFound | if the design has no such combination |
| const ExperimentalDesign::SampleSection & getSampleSection | ( | ) | const |
| std::map< std::string, unsigned > getSampleToConditionMapping | ( | ) | const |
return Sample name to condition mapping (a condition is a unique combination of all columns in the sample section, except for replicates. Numbering of conditions is alphabetical due to map. Keyed by sample NAME, like getSampleToPrefractionationMapping(); it previously used the stringified zero-based sample row index, which only matched designs OpenMS inferred itself.
| std::map< std::string, unsigned > getSampleToPrefractionationMapping | ( | ) | const |
uses getUniqueSampleRowToSampleMapping to get the reversed map mapping sample ID to a real unique sample. Keyed by sample NAME (the Sample column value), not by the zero-based sample row index.
| std::map< std::vector< std::string >, std::set< std::string > > getUniqueSampleRowToSampleMapping | ( | ) | const |
returns a map from a sample section row to sample id for clustering duplicate sample rows (e.g. to find all fractions of the same "sample"). Rows are compared over ALL factors except Sample – replicate columns included – so samples are grouped only when their metadata is completely identical. This is the weaker of the two grouping rules; see getConditionToSampleMapping() for the other one.
| bool isFractionated | ( | ) | const |
Whether the design is fractionated.
|
private |
|
private |
Generic Mapper (Path, Label) -> f(row)
| bool sameNrOfMSFilesPerFraction | ( | ) | const |
|
private |
| void setMSFileSection | ( | const MSFileSection & | msfile_section | ) |
| void setSampleSection | ( | const SampleSection & | sample_section | ) |
|
private |
|
private |
|
private |