bioRxiv ScienceSearch

Biology subjects

Snoek, L.

Publications and source records attributed to Snoek, L..

2 recordsLinked to original sources

How to control for confounds in decoding analyses of neuroimaging data

Over the past decade, multivariate pattern analyses and especially decoding analyses have become a popular alternative to traditional mass-univariate analyses in neuroimaging research. However, a fundamental limitation of decoding analyses is that the source of information driving the decoder is ambiguous, which becomes problematic when the to-be-decoded variable is confounded by variables that are not of primary interest. In this study, we use a comprehensive set of simulations and analyses of empirical data to evaluate two techniques that were previously proposed and used to control for confounding variables in decoding analyses: counterbalancing and confound regression. For our empirical analyses, we attempt to decode gender from structural MRI data when controlling for the confound brain size. We show that both methods introduce strong biases in decoding performance: counterbalancing leads to better performance than expected (i.e., positive bias), which we show in our simulations is due to the subsampling process that tends to remove samples that are hard to classify; confound regression, on the other hand, leads to worse performance than expected (i.e., negative bias), even resulting in significant below-chance performance in some scenarios. In our simulations, we show that below-chance accuracy can be predicted by the variance of the distribution of correlations between the features and the target. Importantly, we show that this negative bias disappears in both the empirical analyses and simulations when the confound regression procedure performed in every fold of the cross-validation routine, yielding plausible model performance. From these results, we conclude that foldwise confound regression is the only method that appropriately controls for confounds, which thus can be used to gain more insight into the exact source(s) of information driving ones decoding analysis.\n\nHIGHLIGHTSO_LIThe interpretation of decoding models is ambiguous when dealing with confounds;\nC_LIO_LIWe evaluate two methods, counterbalancing and confound regression, in their ability to control for confounds;\nC_LIO_LIWe find that counterbalancing leads to positive bias because it removes hard-to-classify samples;\nC_LIO_LIWe find that confound regression leads to negative bias, because it yields data with less signal than expected by chance;\nC_LIO_LIOur simulations demonstrate a tight relationship between model performance in decoding analyses and the sample distribution of the correlation coefficient;\nC_LIO_LIWe show that the negative bias observed in confound regression can be remedied by cross-validating the confound regression procedure;\nC_LI

neuroscience

Porcupine: a visual pipeline tool for neuroimaging analysis

The field of neuroimaging is rapidly adopting a more reproducible approach to data acquisition and analysis. Data structures and formats are being standardised and data analyses are getting more automated. However, as data analysis becomes more complicated, researchers often have to write longer analysis scripts, spanning different tools across multiple programming languages. This makes it more difficult to share or recreate code, reducing the reproducibility of the analysis. We present a tool, Porcupine, that constructs ones analysis visually and automatically produces analysis code. The graphical representation improves understanding of the performed analysis, while retaining the flexibility of modifying the produced code manually to custom needs. Not only does Porcupine produce the analysis code, it also creates a shareable environment for running the code, in the form of a Docker image. Together, this forms a reproducible way of constructing, visualising and sharing ones analysis. Currently, Porcupine links to Nipype functionalities, which in turn accesses most standard neuroimaging analysis tools. With Porcupine, we bridge the gap between a conceptual and an implementational level of analysis and thus create reproducible and shareable science. We give the researcher a better oversight of their pipeline, both while developing and communicating their work. This will reduce the threshold at which less expert users can generate reusable pipelines. We provide a wide range of examples and documentation, as well as installer files for all platforms on our website: https://timvanmourik.github.io/Porcupine. Porcupine is free, open source, andreleased under the GNU General Public License v3.0.\n\nAuthor SummaryThe neuroimaging community is fervently debating that its reproducibility and transparency should be improved, but it is a challenging problem as to how to accomplish this. We here propose a tool, Porcupine, to aid in this process by more easily creating shareable workflows for analysing neuroimaging data. The conceptual understanding of a pipeline is improved by means of the graphical interface, and it automatically produces the code to perform the analysis and to create a sharing environment. This retains full flexibility to modify the script afterwards but in principle produces readily executable code for an end-to-end analysis. Porcupine currently links to all Nipype functionality, but is designed to be extendable to other workflow packages in neuroimaging and beyond. Porcupine is free and is released under the GNU General Public License.

neuroscience