References
Aiken, L. S., & West, S. G. (1991). Multiple regression: Testing
and interpreting interactions. Sage.
Albers, C., & Lakens, D. (2018). When power analyses based on pilot
data are biased: Inaccurate effect size estimators and follow-up bias.
Journal of Experimental Social Psychology, 74,
187–195. https://doi.org/10.1016/j.jesp.2017.09.004
Anscombe, F. J. count. (1973). Graphs in statistical analysis. The
American Statistician, 27(1), 17–21. https://doi.org/10.1080/00031305.1973.10478966
Arel-Bundock, V. (2022). modelsummary: Data
and model summaries in R. Journal of Statistical
Software, 103(1), 1–23. https://doi.org/10.18637/jss.v103.i01
Beerendonk, L., Mejías, J. F., Nuiten, S. A., Gee, J. W. de, Fahrenfort,
J. J., & Gaal, S. van. (2024). A disinhibitory circuit mechanism
explains a general principle of peak performance during mid-level
arousal. Proceedings of the National Academy of Sciences,
121(5). https://doi.org/10.1073/pnas.2312898121
Begley, C. G., & Ellis, L. M. (2012). Raise standards for
preclinical cancer research. Nature, 483(7391),
531–533. https://doi.org/10.1038/483531a
Bem, D. J. (2011). Feeling the future: Experimental evidence for
anomalous retroactive influences on cognition and affect. Journal of
Personality and Social Psychology, 100(3), 407–425. https://doi.org/10.1037/a0021524
Ben-Shachar, M. S., Lüdecke, D., & Makowski, D. (2020). effectsize: Estimation of effect size indices and
standardized parameters. Journal of Open Source Software,
5(56), 2815. https://doi.org/10.21105/joss.02815
Bland, J. M., & Altman, D. G. (2011). Comparisons against baseline
within randomised groups are often used and can be highly misleading.
Trials, 12, 264. https://doi.org/10.1186/1745-6215-12-264
Borm, G. F., Fransen, J., & Lemmens, W. A. (2007). A simple sample
size formula for analysis of covariance in randomized clinical trials.
Journal of Clinical Epidemiology, 60(12), 1234–1238.
https://doi.org/10.1016/j.jclinepi.2007.02.006
Bortz, J., & Schuster, C. (2010). Statistik für human- und
sozialwissenschaftler (7th ed.). Springer.
Box, G. E. P. (1966). Use and abuse of regression.
Technometrics, 8(4), 625. https://doi.org/10.2307/1266635
Box, G. E. P. (1976). Science and statistics. Journal of the
American Statistical Association, 71(356), 791–799.
Brown, N. J. L., & Heathers, J. A. J. (2016). The GRIM test: A
simple technique detects numerous anomalies in the reporting of results
in psychology. Social Psychological and Personality Science,
8(4), 363–369. https://doi.org/10.1177/1948550616673876
Cairo, A. (2012). The functional art: An introduction to information
graphics and visualization. New Riders.
Cleveland, W. S., & McGill, R. (1984). Graphical perception: Theory,
experimentation, and application to the development of graphical
methods. Journal of the American Statistical Association,
79(387), 531–554. https://doi.org/10.1080/01621459.1984.10478080
Cohen, J. (1968). Multiple regression as a general data-analytic system.
Psychological Bulletin, 70(6p1), 426.
Cohen, J. (1988). Statistical power analysis for the behavioural
sciences (2nd ed.). Lawrence Erlbaum Associates.
Cohen, J. (1990). Things i have learned (so far). American
Psychologist, 45(12), 1304–1312. https://doi.org/10.1037/0003-066X.45.12.1304
Cohen, J., Cohen, P., West, S. G., & Aiken, L. S. (2003).
Applied multiple regression/correlation analysis for the behavioral
sciences (3rd ed.). Lawrence Erlbaum.
Corti, L. (2020). Managing and sharing research data: A guide to
good practice (V. V. den Eynden, L. Bishop, M. Woollard, M. Haaker,
& S. Summers, Eds.; 2nd edition). SAGE.
Cumming, G., & Calin-Jageman, R. (2017). Introduction to the new
statistics: Estimation, open science, and beyond. Routledge.
Cumming, Geoff. (2012). Understanding the new statistics: Effect
sizes, confidence intervals, and meta-analysis. Routledge,.
Cumming, G., & Finch, S. (2005). Inference by eye: Confidence
intervals and how to read pictures of data. American
Psychologist, 60(2), 170–180. https://doi.org/10.1037/0003-066X.60.2.170
Curran, P. G. (2016). Methods for the detection of carelessly invalid
responses in survey data. Journal of Experimental Social
Psychology, 66, 4–19. https://doi.org/10.1016/j.jesp.2015.07.006
Döring, N. (2023). Forschungsmethoden und evaluation: Für human- und
sozialwissenschaftler. Springer.
Eid, M., Gollwitzer, M., & Schmitt, M. (2017). Statistik und
forschungsmethoden.
Faul, F., Erdfelder, E., G., L. A., & Buchner, A. (2007). G*power 3:
A flexible statistical power analysis program for the social,
behavioral, and biomedical sciences. Behavior Research Methods,
39, 175–191.
Field, A. (2026). Discovering statistics using
R (2nd edition). Sage.
Fisher, R. A. (1922). On the mathematical foundations of theoretical
statistics. Philosophical Transactions of the Royal Society of
London. Series A, 222, 309–368. https://doi.org/10.1098/rsta.1922.0009
Fisher, R. A. (1925). Statistical methods for research workers.
Oliver; Boyd.
Fisher, R. A. (1938). Presidential address to the first
Indian statistical congress. Sankhyā,
4(1), 14–17.
Fox, J. (2016). Applied regression analysis and generalized linear
models (3rd ed.). Sage.
Galilei, G. (2016). Sidereus nuncius, or the sidereal
messenger. University of Chicago Press. (Original work published
1610)
Galton, F. (1886). Regression towards mediocrity in hereditary stature.
The Journal of the Anthropological Institute of Great Britain and
Ireland, 15, 246–263. https://doi.org/10.2307/2841583
Gelman, A. (2007). Scaling regression inputs by dividing by two standard
deviations. Statistics in Medicine, 27(15), 2865–2873.
https://doi.org/10.1002/sim.3107
Gelman, A., Hill, J., & Vehtari, A. (2020). Regression and other
stories. Cambridge University Press. https://doi.org/10.1017/9781139161879
Gelman, A., & Loken, E. (2014). The statistical crisis in science.
American Scientist, 102(6), 460. https://doi.org/10.1511/2014.111.460
Gelman, A., & Stern, H. (2006). The difference between significant
and not significant is not itself statistically significant.
American Statistician, 60(4), 328–331.
Greenhouse, S. W., & Geisser, S. (1959). On methods in the analysis
of profile data. Psychometrika, 24(2), 95–112. https://doi.org/10.1007/BF02289823
Herndon, T., Ash, M., & Pollin, R. (2013). Does high public debt
consistently stifle economic growth? A critique of reinhart and rogoff.
Cambridge Journal of Economics, 38(2), 257–279. https://doi.org/10.1093/cje/bet075
Hunsley, J., & Meyer, G. J. (2003). The incremental validity of
psychological testing and assessment: Conceptual, methodological, and
statistical issues. Psychological Assessment, 15(4),
446–455. https://doi.org/10.1037/1040-3590.15.4.446
Johnson, J. W., & LeBreton, J. M. (2004). History and use of
relative importance indices in organizational research.
Organizational Research Methods, 7(3), 238–257. https://doi.org/10.1177/1094428104266510
Lakens, D. (2013). Calculating and reporting effect sizes to facilitate
cumulative science: A practical primer for t-tests and ANOVAs.
Frontiers in Psychology, 4, 863.
Lakens, D. (2022). Sample size justification. Collabra:
Psychology, 8(1), 33267. https://doi.org/10.1525/collabra.33267
Lehr, R. (1992). Sixteen s-squared over d-squared: A relation for crude
sample size estimates. Statistics in Medicine, 11(8),
1099–1102. https://doi.org/10.1002/sim.4780110811
Lindeløv, J. K. (2019). Common statistical tests are linear models
(or: How to teach stats). https://lindeloev.github.io/tests-as-linear/.
Long, J. S., & Ervin, L. H. (2000). Using heteroscedasticity
consistent standard errors in the linear regression model. The
American Statistician, 54(3), 217–224. https://doi.org/10.1080/00031305.2000.10474549
Maslow, A. (1962). Toward a psychology of being. D Van
Nostrand. https://doi.org/10.1037/10793-000
Matejka, J., & Fitzmaurice, G. (2017). Same stats, different graphs:
Generating datasets with varied appearance and identical statistics
through simulated annealing. Proceedings of the 2017 CHI Conference
on Human Factors in Computing Systems, 1290–1294. https://doi.org/10.1145/3025453.3025912
Maxwell, S. E., Delaney, H. D., & Kelley, K. (2018). Designing
experiments and analyzing data: A model comparison perspective (3rd
ed.). Routledge. https://doi.org/10.4324/9781315642956
Meade, A. W., & Craig, S. B. (2012). Identifying careless responses
in survey data. Psychological Methods, 17(3), 437–455.
https://doi.org/10.1037/a0028085
Meyer, G. J., Finn, S. E., Eyde, L. D., Kay, G. G., Moreland, K. L.,
Dies, R. R., Eisman, E. J., Kubiszyn, T. W., & Reed, G. M. (2001).
Psychological testing and psychological assessment: A review of evidence
and issues. American Psychologist, 56(2), 128–165. https://doi.org/10.1037/0003-066x.56.2.128
Miller, G. A., & Chapman, J. P. (2001). Misunderstanding analysis of
covariance. Journal of Abnormal Psychology, 110(1),
40–48. https://doi.org/10.1037/0021-843X.110.1.40
Moosbrugger, H., & Kelava, A. (Eds.). (2020). Testtheorie und
fragebogenkonstruktion. Springer Berlin Heidelberg. https://doi.org/10.1007/978-3-662-61532-4
Morris, S. B. (2008). Estimating effect sizes from
pretest-posttest-control group designs. Organizational Research
Methods, 11(2), 364–386. https://doi.org/10.1177/1094428106291059
Morris, T. P., White, I. R., & Crowther, M. J. (2019). Using
simulation studies to evaluate statistical methods. Statistics in
Medicine, 38(11), 2074–2102. https://doi.org/10.1002/sim.8086
Munafò, M. R. et al. (2017). A manifesto for
reproducible science. Nature Human Behaviour, 1, 0021.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T.
(2018). The preregistration revolution. Proceedings of the National
Academy of Sciences, 115(11), 2600–2606. https://doi.org/10.1073/pnas.1708274114
Open Science Collaboration. (2015). Estimating the reproducibility of
psychological science. Science, 349(6251). https://doi.org/10.1126/science.aac4716
Orben, A., & Przybylski, A. K. (2019). The association between
adolescent well-being and digital technology use. Nature Human
Behaviour, 3(2), 173–182. https://doi.org/10.1038/s41562-018-0506-1
Popper, K. R. (2007). Conjectures and refutations: The growth of
scientific knowledge (Repr.). Routledge.
R Core Team. (2026). R: A language and environment for statistical
computing. R Foundation for Statistical Computing. https://doi.org/10.32614/R.manuals
Reinhart, C. M., & Rogoff, K. S. (2010). Growth in a time of debt.
American Economic Review, 100(2), 573–578. https://doi.org/10.1257/aer.100.2.573
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022).
Revisiting meta-analytic estimates of validity in personnel selection:
Addressing systematic overcorrection for restriction of range.
Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994
Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of
selection methods in personnel psychology: Practical and theoretical
implications of 85 years of research findings. Psychological
Bulletin, 124(2), 262–274. https://doi.org/10.1037/0033-2909.124.2.262
Sedlmeier, P., & Gigerenzer, G. (1989). Do studies of statistical
power have an effect on the power of studies? Psychological
Bulletin, 105(2), 309–316.
Senn, S. (2006). Change from baseline and analysis of covariance
revisited. Statistics in Medicine, 25(24), 4334–4344.
https://doi.org/10.1002/sim.2682
Shmueli, G. (2010). To explain or to predict? Statistical
Science, 25(3), 289–310. https://doi.org/10.1214/10-sts330
Shu, L. L., Mazar, N., Gino, F., Ariely, D., & Bazerman, M. H.
(2012). RETRACTED: Signing at the beginning makes ethics salient and
decreases dishonest self-reports in comparison to signing at the end.
Proceedings of the National Academy of Sciences,
109(38), 15197–15200. https://doi.org/10.1073/pnas.1209746109
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011).
False-positive psychology: Undisclosed flexibility in data collection
and analysis allows presenting anything as significant.
Psychological Science, 22(11), 1359–1366.
Tufte, E. R. (1983). The visual display of quantitative
information. Graphics press.
Tukey, J. W. (1962). The future of data analysis. The Annals of
Mathematical Statistics, 33(1), 1–67. https://doi.org/10.1214/aoms/1177704711
Tukey, J. W. (1977). Exploratory data analysis. Reading, Mass.
Van Breukelen, G. J. P. (2006). ANCOVA versus change from
baseline had more power in randomized studies and more bias in
nonrandomized studies. Journal of Clinical Epidemiology,
59(9), 920–925. https://doi.org/10.1016/j.jclinepi.2006.02.007
Vickers, A. J., & Altman, D. G. (2001). Analysing controlled trials
with baseline and follow up measurements. BMJ,
323(7321), 1123–1124. https://doi.org/10.1136/bmj.323.7321.1123
Wagenmakers, E.-J., Marsman, M., Jamil, T., Ly, A., Verhagen, J., Love,
J., Selker, R., Gronau, Q. F., Šmíra, M., Epskamp, S., Matzke, D.,
Rouder, J. N., & Morey, R. D. (2017). Bayesian inference for
psychology. Part i: Theoretical advantages and practical ramifications.
Psychonomic Bulletin &Amp; Review, 25(1), 35–57.
https://doi.org/10.3758/s13423-017-1343-3
Westfall, J., & Yarkoni, T. (2016). Statistically controlling for
confounding constructs is harder than you think. PLoS ONE,
11(3), e0152719. https://doi.org/10.1371/journal.pone.0152719
Wickham, H. (2010). A layered grammar of graphics. Journal of
Computational and Graphical Statistics, 19(1), 3–28. https://doi.org/10.1198/jcgs.2009.07098
Wickham, H. (2014). Tidy data. Journal of Statistical Software,
59. https://doi.org/10.18637/jss.v059.i10
Wickham, H. (2016). ggplot2: Elegant graphics for data analysis
(2nd ed.). Springer. https://doi.org/10.1007/978-3-319-24277-4
Wilkinson, L. (2012). The grammar of graphics. In Handbook of
computational statistics. Springer. https://doi.org/10.1007/978-3-642-21551-3_13
Wilkinson, M. D., Dumontier, M., Aalbersberg, Ij. J., Appleton, G.,
Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., Silva Santos, L. B.
da, Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M.,
Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., … Mons, B.
(2016). The FAIR guiding principles for scientific data management and
stewardship. Scientific Data, 3(1). https://doi.org/10.1038/sdata.2016.18
Williams, M. N., Grajales, C. A. G., & Kurkiewicz, D. (2013).
Assumptions of multiple regression: Correcting two misconceptions.
Practical Assessment, Research, and Evaluation,
18(11). https://doi.org/10.7275/55hn-wk47