|
| void | computeMedians_ (SeqToList &rt_data, SeqToValue &medians, bool sorted=false) |
| | Compute the median retention time for each peptide sequence.
|
| |
| bool | getRetentionTimes_ (const PeptideIdentificationList &peptides, SeqToList &rt_data) |
| | Collect retention time data from peptide IDs.
|
| |
| bool | getRetentionTimes_ (const IdentificationData &id_data, SeqToList &rt_data) |
| | Collect retention time data from spectrum matches.
|
| |
| bool | getRetentionTimes_ (const IsFCMap auto &features, SeqToList &rt_data) |
| | Collect retention time data from peptide IDs contained in feature maps or consensus maps.
|
| |
| void | computeTransformations_ (std::vector< SeqToList > &rt_data, std::vector< TransformationDescription > &transforms, bool sorted=false, bool verbose=true) |
| | Compute retention time transformations from RT data grouped by peptide sequence.
|
| |
| Int | selectReference_ (const std::vector< SeqToList > &rt_data) const |
| | Choose the input map that shares the most identified sequences with every other map as the reference.
|
| |
| void | alignToAutoReference_ (std::vector< SeqToList > &rt_data, std::vector< TransformationDescription > &transforms, bool sorted) |
| | Compute RT transformations without a given reference, as parameter auto_reference asks.
|
| |
| void | alignToInput_ (std::vector< SeqToList > &rt_data, Size index, std::vector< TransformationDescription > &transforms, bool sorted, bool verbose) |
| | Compute RT transformations with one of the input maps as the reference.
|
| |
| void | checkParameters_ (const Size runs) |
| | Check that parameter values are valid.
|
| |
| void | getReference_ () |
| | Get reference retention times.
|
| |
| IdentificationData::ScoreTypeRef | handleIdDataScoreType_ (const IdentificationData &id_data) |
| | Helper function to find/define the score type for processing IdentificationData.
|
| |
| const PeptideHit * | getBestScoringHit (const std::vector< PeptideHit > &hits, const bool is_higher_score_better) |
| | Get the best-scoring PeptideHit from a list of hits.
|
| |
Protected Member Functions inherited from DefaultParamHandler |
| virtual void | updateMembers_ () |
| | This method is used to update extra member variables at the end of the setParameters() method.
|
| |
| void | defaultsToParam_ () |
| | Updates the parameters after the defaults have been set in the constructor.
|
| |
|
| Int | reference_index_ |
| | Index of input file to use as reference (if any)
|
| |
| SeqToValue | reference_ |
| | Reference retention times (per peptide sequence)
|
| |
| Size | min_run_occur_ |
| | Minimum number of runs a peptide must occur in.
|
| |
| bool | use_feature_rt_ {} |
| | Use feature RT instead of RT from best peptide ID in the feature?
|
| |
| bool | use_adducts_ {} |
| | Consider differently adducted IDs as different?
|
| |
| bool | consensus_reference_ {} |
| | Without a given reference, align to a consensus of all maps (instead of the map that shares the most IDs with the others)?
|
| |
| Size | auto_reference_min_points_ {} |
| | Number of alignment points that the automatically chosen reference map should provide for every other map (else a consensus is tried)
|
| |
| double | min_score_ |
| | Minimum score to reach for a peptide to be considered.
|
| |
| bool | score_cutoff_ {} |
| | Actually use the above defined score_cutoff? Needed since it is hard to define a non-cutting score for a user.
|
| |
| std::string | score_type_ |
| | Score type to use for filtering.
|
| |
| bool(* | better_ )(double, double) = [](double, double) {return true;} |
| | Score better?
|
| |
Protected Attributes inherited from DefaultParamHandler |
| Param | param_ |
| | Container for current parameters.
|
| |
| Param | defaults_ |
| | Container for default parameters. This member should be filled in the constructor of derived classes!
|
| |
| std::vector< std::string > | subsections_ |
| | Container for registered subsections. This member should be filled in the constructor of derived classes!
|
| |
| std::string | error_name_ |
| | Name that is displayed in error messages during the parameter checking.
|
| |
| bool | check_defaults_ |
| | If this member is set to false no checking if parameters in done;.
|
| |
| bool | warn_empty_defaults_ |
| | If this member is set to false no warning is emitted when defaults are empty;.
|
| |
| LogType | type_ |
| |
| time_t | last_invoke_ |
| |
| ProgressLoggerImpl * | current_logger_ |
| |
A map alignment algorithm based on peptide identifications from MS2 spectra.
PeptideIdentification instances are grouped by sequence of the respective best-scoring PeptideHit and retention time data is collected (PeptideIdentification::getRT()). ID groups with the same sequence in different maps represent points of correspondence between the maps and form the basis of the alignment. Only the best PSM per spectrum is considered as the correct identification.
Each map is aligned to a reference retention time scale. This time scale can come from a separate reference (setReference()) or from one of the input maps (reference_index in align()). If neither is given, parameter auto_reference decides: by default ("best_run"), the input map that shares the most identified sequences with every other map becomes the reference, i.e. the map whose smallest number of shared sequences with any other map is largest (on ties, the map with the most identified sequences), so that every map can be aligned to it. With "consensus", the time scale is computed as a consensus of the input maps (median retention times over all maps of the ID groups). A consensus is also used if no map shares at least two sequences with every other map. If the chosen map leaves other maps with fewer than auto_reference_min_points alignment points (after removing outliers, see max_rt_shift), and one of them shares at least that many sequences with other maps, every map and a consensus are tried as the reference, and the choice that gives the most maps at least that many points is used (on ties, the first choice is kept, and a map is preferred over a consensus). A consensus favors none of the maps, but every map contributes to the consensus it is aligned to, so larger shifts between maps are only partly corrected. The maps are then aligned to this scale as follows:
The median retention time of each ID group in a map is mapped to the reference retention time of this group. Cubic spline smoothing is used to convert this mapping to a smooth function. Retention times in the map are transformed to the consensus scale by applying this function.
Parameters of this class are:
| Name | Type | Default | Restrictions | Description |
| score_type |
string | |
| Name of the score type to use for ranking and filtering (.oms input only). If left empty, a score type is picked automatically. |
| score_cutoff |
string | false |
true, false | Use only IDs above a score cut-off (parameter 'min_score') for alignment? |
| min_score |
float | 0.05 |
| If 'score_cutoff' is 'true': Minimum score for an ID to be considered. Unless you have very few runs or identifications, increase this value to focus on more informative peptides. |
| min_run_occur |
int | 2 |
min: 2 | Minimum number of runs (incl. reference, if any) in which a peptide must occur to be used for the alignment. Unless you have very few runs or identifications, increase this value to focus on more informative peptides. |
| max_rt_shift |
float | 0.5 |
min: 0.0 | Maximum realistic RT difference for a peptide (median per run vs. reference). Peptides with higher shifts (outliers) are not used to compute the alignment. If 0, no limit (disable filter); if > 1, the final value in seconds; if <= 1, taken as a fraction of the range of the reference RT scale. |
| use_unassigned_peptides |
string | true |
true, false | Should unassigned peptide identifications be used when computing an alignment of feature or consensus maps? If 'false', only peptide IDs assigned to features will be used. |
| use_feature_rt |
string | false |
true, false | When aligning feature or consensus maps, don't use the retention time of a peptide identification directly; instead, use the retention time of the centroid of the feature (apex of the elution profile) that the peptide was matched to. If different identifications are matched to one feature, only the peptide closest to the centroid in RT is used. Precludes 'use_unassigned_peptides'. |
| use_adducts |
string | true |
true, false | If IDs contain adducts, treat differently adducted variants of the same molecule as different. |
| auto_reference |
string | best_run |
best_run, consensus | Reference to align to if none is given (neither a reference file nor an input index): 'best_run' - the input that shares the most identified sequences with every other input (on ties, the one with the most identified sequences). A consensus is used instead if no input shares at least two sequences with every other input. If the chosen input leaves other inputs with too few alignment points, other inputs and a consensus are tried as well (see 'auto_reference_min_points'). 'consensus' - median RTs per sequence over all inputs. A consensus favors none of the inputs, but only partly corrects larger RT shifts, because every input contributes to the consensus it is aligned to. |
| auto_reference_min_points |
int | 11 |
min: 0 | If 'auto_reference' is 'best_run': number of alignment points (after removing outliers, see 'max_rt_shift') that the reference should provide for every other input. If the chosen input leaves inputs with fewer points, and one of them shares at least this many sequences with other inputs, every input and a consensus of all inputs are tried as the reference. The choice that gives the most inputs at least this many points is used (the reference counts); on ties, the first choice is kept, and an input is preferred over a consensus. The default is the smallest number of points to which ProteomicsLFQ and MS1LabeledWorkflow fit an RT model. 0 disables the check. |
Note:
- If a section name is documented, the documentation is displayed as tooltip.
- Advanced parameter names are italic.
Compute RT transformations without a given reference, as parameter auto_reference asks.
With "best_run", aligns to the map chosen by selectReference_() and sets reference_index_ accordingly. If that leaves maps with fewer than auto_reference_min_points alignment points, and one of them shares at least that many sequences with other maps, every map and a consensus of all maps are tried as the reference. The choice that gives the most maps at least that many points (the reference counts) is used; on ties, the first choice is kept, then the map with the largest smallest number of points among those maps is preferred, and a map is preferred over a consensus.
- Parameters
-
| [in,out] | rt_data | Lists of RT values for diff. peptide sequences, per input map (input, will be sorted) |
| [out] | transforms | Resulting transformations, per input map (output) |
| [in] | sorted | Are RT lists already sorted? |
| bool getRetentionTimes_ |
( |
const IsFCMap auto & |
features, |
|
|
SeqToList & |
rt_data |
|
) |
| |
|
inlineprotected |
Collect retention time data from peptide IDs contained in feature maps or consensus maps.
The following global flags (mutually exclusive) influence the processing:
Depending on use_unassigned_peptides, unassigned peptide IDs are used in addition to IDs annotated to features.
Depending on use_feature_rt, feature retention times are used instead of peptide retention times. Depending on score_cutoff and min_score, only peptide IDs with minimum score X are used. Higher score better is determined from the first PeptideID encountered. Make sure they are the same. This param is useless with use_feature_rt yet.
- Parameters
-
| [in] | features | Input features for RT data |
| [out] | rt_data | Lists of RT values for diff. peptide sequences (output) |
- Returns
- Are the RTs already sorted? (Here: true)
References PeptideHit::getScore(), PeptideHit::getSequence(), and AASequence::toString().