The science behind the suite

Every method, named,
with the paper behind it.

If a number from one of these tools is going into a figure, you should be able to write its method down. This page is the list: what each tool computes, and where each of those methods comes from.

Three rules govern this page. Nothing is listed that is not in the code: the methods below were read off the engines themselves, not off a feature list. Every reference was resolved before being written here, so each link goes to the actual paper. And a method we have not implemented does not appear, however standard it is.

How these implementations are checked against their originals is a separate question, answered on the validation page. What to write in your own methods section is on the citation page.

Graphs & statistics

TwistPlot

The tests run in R on our side, and the figure shows what was run: the test name, the correction, the n, and the exact p-value rather than a row of stars. When a design does not meet the assumptions of the test you picked, the tool says so instead of quietly switching.

What it computes

  • Group comparisons: Welch and Student t-tests, one- and two-way ANOVA, and their rank-based counterparts
  • Post-hoc comparisons: Tukey for all pairs, Dunnett against a single control, Dunn after a rank-based test
  • Multiplicity: Holm, Bonferroni and Benjamini-Hochberg
  • Assumption checks: Shapiro-Wilk or D'Agostino-Pearson for normality, Levene centred on the median for equality of variances
  • Survival: Kaplan-Meier curves, the log-rank test overall and pairwise, and hazard ratios from Cox regression
  • ROC curves, with the area under the curve for each group
  • Dose-response: four-parameter logistic fits with IC50 and EC50

References

See TwistPlot in action →

ELISA

Sorbet

Standard curves are fitted by nonlinear least squares with the Levenberg-Marquardt algorithm, the same one the R original used, and the back-calculated concentrations carry the curve's own uncertainty rather than a straight-line approximation of it.

What it computes

  • Three-, four- and five-parameter logistic standard curves
  • Levenberg-Marquardt nonlinear least squares, with the fit diagnostics that go with it
  • Back-calculation of unknowns, with dilution factors applied
  • Limit of detection and limit of quantification from the blanks and the low standards

References

  • Marquardt, D. W. (1963). An algorithm for least-squares estimation of nonlinear parameters. Journal of the SIAM 11, 431-441. doi:10.1137/0111030
  • Gottschalk, P. G. & Dunn, J. R. (2005). The five-parameter logistic: a characterization and comparison with the four-parameter logistic. Analytical Biochemistry 343, 54-65. doi:10.1016/j.ab.2005.04.035

See Sorbet in action →

qPCR & primer design

CtExplorer

Relative quantification is efficiency-corrected by default, because assuming every assay doubles perfectly is the most common way a qPCR figure goes wrong. Primer design uses published thermodynamics rather than a rule of thumb.

What it computes

  • Relative expression by the Pfaffl model, with measured amplification efficiencies
  • The 2^-DDCt method where efficiencies are assumed equal, side by side with the above
  • Reference gene stability, and normalisation against the geometric mean of several genes
  • Melting curve analysis and per-plate quality control
  • Statistics on the normalised values, with correction for multiple genes
  • Primer design: melting temperatures from nearest-neighbour thermodynamics with a salt correction, hairpin and dimer scoring on the Primer3 alignment model

References

  • Pfaffl, M. W. (2001). A new mathematical model for relative quantification in real-time RT-PCR. Nucleic Acids Research 29, e45. doi:10.1093/nar/29.9.e45
  • Livak, K. J. & Schmittgen, T. D. (2001). Analysis of relative gene expression data using real-time quantitative PCR and the 2^-DDCt method. Methods 25, 402-408. doi:10.1006/meth.2001.1262
  • Vandesompele, J. et al. (2002). Accurate normalization of real-time quantitative RT-PCR data by geometric averaging of multiple internal control genes. Genome Biology 3, research0034. doi:10.1186/gb-2002-3-7-research0034
  • SantaLucia, J. (1998). A unified view of polymer, dumbbell and oligonucleotide DNA nearest-neighbor thermodynamics. PNAS 95, 1460-1465. doi:10.1073/pnas.95.4.1460
  • von Ahsen, N., Wittwer, C. T. & Schutz, E. (2001). Oligonucleotide melting temperatures under PCR conditions. Clinical Chemistry 47, 1956-1961. doi:10.1093/clinchem/47.11.1956
  • Untergasser, A. et al. (2012). Primer3: new capabilities and interfaces. Nucleic Acids Research 40, e115. doi:10.1093/nar/gks596
  • Welch, B. L. (1947). The generalization of Student's problem when several different population variances are involved. Biometrika 34, 28-35. doi:10.1093/biomet/34.1-2.28
  • Mann, H. B. & Whitney, D. R. (1947). On a test of whether one of two random variables is stochastically larger than the other. Annals of Mathematical Statistics 18, 50-60. doi:10.1214/aoms/1177730491
  • Kruskal, W. H. & Wallis, W. A. (1952). Use of ranks in one-criterion variance analysis. JASA 47, 583-621. doi:10.1080/01621459.1952.10483441
  • Tukey, J. W. (1949). Comparing individual means in the analysis of variance. Biometrics 5, 99-114. doi:10.2307/3001913
  • Benjamini, Y. & Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing. JRSS B 57, 289-300. doi:10.1111/j.2517-6161.1995.tb02031.x

See CtExplorer in action →

Spectral flow panels

CytoMix

Panel scoring is spectral: it works from the full emission signatures of the dyes on the instrument you selected, not from a table of "avoid these pairs". The FCS files you acquire are read in the browser and never leave your computer.

What it computes

  • Similarity between full emission spectra on a chosen instrument configuration, and the spillover that follows
  • Panel scoring on overlap, tandem dyes, antigen density and co-expression
  • Master mix and titration volumes, and the FMO controls that go with the panel
  • FCS 3.1 file reading, with logicle and arcsinh display transforms
  • Exploratory clustering on the events: nearest-neighbour graph, Leiden communities, UMAP embedding

References

  • Spidlen, J. et al. (2010). Data file standard for flow cytometry, version FCS 3.1. Cytometry Part A 77A, 97-100. doi:10.1002/cyto.a.20825
  • Parks, D. R., Roederer, M. & Moore, W. A. (2006). A new "logicle" display method avoids deceptive effects of logarithmic scaling for low signals and compensated data. Cytometry Part A 69A, 541-551. doi:10.1002/cyto.a.20258
  • Novo, D., Gregori, G. & Rajwa, B. (2013). Generalized unmixing model for multispectral flow cytometry utilizing nonsquare compensation matrices. Cytometry Part A 83A, 508-520. doi:10.1002/cyto.a.22272
  • Maecker, H. T., Frey, T., Nomura, L. E. & Trotter, J. (2004). Selecting fluorochrome conjugates for maximum sensitivity. Cytometry Part A 62A, 169-173. doi:10.1002/cyto.a.20092
  • Traag, V. A., Waltman, L. & van Eck, N. J. (2019). From Louvain to Leiden: guaranteeing well-connected communities. Scientific Reports 9, 5233. doi:10.1038/s41598-019-41695-z
  • McInnes, L., Healy, J. & Melville, J. (2018). UMAP: uniform manifold approximation and projection for dimension reduction. arXiv:1802.03426. arXiv

See CytoMix in action →

FCS toolbox

FlowFusion

A toolbox rather than a pipeline: it reads what your instrument and your analysis software already wrote, including FlowJo workspaces and SpectroFlo exports, and gives back files of the same kind.

What it computes

  • FCS 3.1 reading, rewriting, merging, renaming and subsampling, including large spectral files
  • FlowJo workspace and SpectroFlo gating read as they are
  • Titration series with the stain index computed per concentration
  • Logicle and biexponential scaling, matched to what the acquisition software displays

References

  • Spidlen, J. et al. (2010). Data file standard for flow cytometry, version FCS 3.1. Cytometry Part A 77A, 97-100. doi:10.1002/cyto.a.20825
  • Parks, D. R., Roederer, M. & Moore, W. A. (2006). A new "logicle" display method avoids deceptive effects of logarithmic scaling for low signals and compensated data. Cytometry Part A 69A, 541-551. doi:10.1002/cyto.a.20258
  • Maecker, H. T., Frey, T., Nomura, L. E. & Trotter, J. (2004). Selecting fluorochrome conjugates for maximum sensitivity. Cytometry Part A 62A, 169-173. doi:10.1002/cyto.a.20092

See FlowFusion in action →

Bulk & single-cell RNA-seq

LavaSEQ

Bulk differential expression runs in R on our side, with the Bioconductor packages the field reads and reviews rather than a reimplementation of them. Quality control, normalisation and the whole single-cell path run in your browser, following the reference implementations step by step, because that is what they are tested against.

What it computes

  • Bulk differential expression with DESeq2 and limma, on the contrasts you define
  • Library-size normalisation and expression filtering as in edgeR, computed in the browser and checked against edgeR's own values
  • Batch correction with sva, when the design calls for it
  • Gene set enrichment with fgsea, against the MSigDB hallmark collection and GO
  • False discovery rate control on every gene-level table
  • Single-cell: quality control, normalisation, highly variable genes, principal components
  • Single-cell: nearest-neighbour graph, Leiden clustering, UMAP embedding, batch integration with Harmony
  • Single-cell: marker genes by rank-sum test, and pseudobulk differential expression where the unit is the sample, not the cell

References

  • Love, M. I., Huber, W. & Anders, S. (2014). Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biology 15, 550. doi:10.1186/s13059-014-0550-8
  • Ritchie, M. E. et al. (2015). limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Research 43, e47. doi:10.1093/nar/gkv007
  • Robinson, M. D., McCarthy, D. J. & Smyth, G. K. (2010). edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics 26, 139-140. doi:10.1093/bioinformatics/btp616
  • Leek, J. T., Johnson, W. E., Parker, H. S., Jaffe, A. E. & Storey, J. D. (2012). The sva package for removing batch effects and other unwanted variation in high-throughput experiments. Bioinformatics 28, 882-883. doi:10.1093/bioinformatics/bts034
  • Johnson, W. E., Li, C. & Rabinovic, A. (2007). Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics 8, 118-127. doi:10.1093/biostatistics/kxj037
  • Korotkevich, G., Sukhov, V. & Sergushichev, A. (2016). Fast gene set enrichment analysis. bioRxiv 060012. doi:10.1101/060012
  • Liberzon, A. et al. (2015). The Molecular Signatures Database hallmark gene set collection. Cell Systems 1, 417-425. doi:10.1016/j.cels.2015.12.004
  • Benjamini, Y. & Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing. JRSS B 57, 289-300. doi:10.1111/j.2517-6161.1995.tb02031.x
  • Wolf, F. A., Angerer, P. & Theis, F. J. (2018). SCANPY: large-scale single-cell gene expression data analysis. Genome Biology 19, 15. doi:10.1186/s13059-017-1382-0
  • Traag, V. A., Waltman, L. & van Eck, N. J. (2019). From Louvain to Leiden: guaranteeing well-connected communities. Scientific Reports 9, 5233. doi:10.1038/s41598-019-41695-z
  • McInnes, L., Healy, J. & Melville, J. (2018). UMAP: uniform manifold approximation and projection for dimension reduction. arXiv:1802.03426. arXiv
  • Korsunsky, I. et al. (2019). Fast, sensitive and accurate integration of single-cell data with Harmony. Nature Methods 16, 1289-1296. doi:10.1038/s41592-019-0619-0
  • Wilcoxon, F. (1945). Individual comparisons by ranking methods. Biometrics Bulletin 1, 80-83. doi:10.2307/3001968

See LavaSEQ in action →

The tools that compute nothing statistical

Three of the nine deliberately have no statistics in them, and it would be dishonest to dress them up: Plaquo draws plate layouts, Dizics plans injections and solutions, and Mirage is a searchable library of illustrations. Their arithmetic is dilutions, volumes and geometry, and it is checked against the same kind of reference values as everything else, but there is no model in them and no paper to cite.

Something missing or misattributed? Write to support@getatlascope.com. A wrong reference on this page is as much a bug as a wrong number in a tool.