Rogue Scholar

PapersBiological Sciences

Weekly Recap (Nov 2024, part 1)

https://doi.org/10.59350/4a13p-nky04

Published November 8, 2024

Author Stephen Turner

This week's recap highlights a new pipeline for metagenome quality assessment and taxonomic annotation (MAGFlow &

NextflowBiological Sciences

Nextflow Summit Barcelona 2024

https://doi.org/10.59350/xkn7n-zep24

Published November 4, 2024

Author Stephen Turner

I just returned from a week in Barcelona where I attended the Nextflow Summit and nf-core hackathon, and I can hardly contain my excitement for the near term future of bioinformatics, computational biology, and open science in general.

PapersBiological Sciences

Weekly Recap (Oct 2024, part 4)

https://doi.org/10.59350/xhh6j-z5b67

Published October 25, 2024

Author Stephen Turner

This week’s recap highlights protein design with RoseTTAFold, surveillance with wastewater sequencing, T2T human genomes, Vitessce for visualization of multimodal spatial single-cell data, and Taxometer for taxonomic classification of metagenomics contigs.

R PythonBiological Sciences

Python for R users

https://doi.org/10.59350/nw1ga-79906

Published October 21, 2024

Author Stephen Turner

A Google search for “R vs Python” returns thousands of hits across sites like Reddit, IBM, Datacamp, Coursera, Kaggle, and many others. A quick Google Trends analysis shows that this search query has grown steadily over the last decade. Any real data scientist would agree that this argument is silly, that the right answer is to use the best tool for the job. What’s “best” isn’t always easy to answer.

PapersBiological Sciences

Weekly Recap (Oct 2024, part 3)

https://doi.org/10.59350/dy8w5-g3p74

Published October 18, 2024

Author Stephen Turner

This week’s recap highlights a new Nextflow workflow for calculating polygenic scores with adjustments for genetic ancestry, a paper demonstrating that whole exome plus imputation on more samples is more powerful than whole genome sequencing for finding more trait associated variants, a new deep-learning-based splice site predictor that improves spliced alignments, a new method for accurate community profiling of large metagenomic datasets, and

PapersAIBiological Sciences

Inciteful+Zotero to find relevant literature

https://doi.org/10.59350/p027a-93g67

Published October 14, 2024

Author Stephen Turner

I am in the middle of writing a review / perspectives paper. One that I’m confident will be exciting once we get it published. Some sections of the review cover subject matter at the outer periphery of my expertise.

PapersBiological Sciences

Weekly Recap (Oct 2024, part 2)

https://doi.org/10.59350/fxnyy-wng24

Published October 11, 2024

Author Stephen Turner

This week’s recap highlights a new method for gene-level alignment of single-cell trajectories, an R package for integrating gene and protein identifiers across biological sequence databases, characterization of SVs across humans and apes, universal prediction of cellular phenotypes, a method to quantify cell state heritability versus plasticity and infer cell state transition with single cell data, and a new AI-driven, natural language-oriented

R TILBiological Sciences

Use nanoparquet instead of readr/CSV

https://doi.org/10.59350/mtjg1-vf107

Published October 8, 2024

Author Stephen Turner

Yesterday I wrote about base R vs. dplyr vs. duckdb for a simple summary analysis. In that post I simulated 100 million rows of a dataset and wrote to disk as CSV. I then benchmarked how long it took to read in and compute a simple grouped mean. One thing I didn’t do here was separate the time it took to read data into memory (for base R and dplyr) versus computing the actual summary.

R TILBiological Sciences

DuckDB vs dplyr vs base R

https://doi.org/10.59350/b4eds-n3h83

Published October 7, 2024

Author Stephen Turner

TL;DR : For a very simple analysis (means by group on 100M rows), duckdb was 125x faster than base R, and 28x faster than readr+dplyr, without having to read data from disk into memory. The duckplyr package wraps DuckDB's analytical query processing techniques in a dplyr-compatible API. Learn more at duckdb.org/docs/api/r and duckplyr.tidyverse.org. I wanted to see for myself what the fuss was about with DuckDB.

PapersBiological Sciences

Weekly Recap (Oct 2024, part 1)

https://doi.org/10.59350/qnb80-c4912

Published October 4, 2024

Author Stephen Turner

This week’s recap highlights a new multispecies codon optimization method, personalized pangenome references with vg, a commentary on the wild west of spike-in normalization, a new pipeline for comprehensive and scalable polygenic scoring across ancestrally diverse populations, a paper showing deep learning / transformer-based methods don’t outperform simple linear models for predicting gene expression after genetic perturbations, and finally, a

R AIPythonBiological Sciences

AI code completion in Positron

https://doi.org/10.59350/bxp4b-neb71

Published October 1, 2024

Author Stephen Turner

TL;DR: Codeium offers a free Copilot-like experience in Positron. You can install it from the Open VSX registry directly within the extensions pane in Positron.

Paired Ends

Weekly Recap (Nov 2024, part 1)

Nextflow Summit Barcelona 2024

Weekly Recap (Oct 2024, part 4)

Python for R users

Weekly Recap (Oct 2024, part 3)

Inciteful+Zotero to find relevant literature

Weekly Recap (Oct 2024, part 2)

Use nanoparquet instead of readr/CSV

DuckDB vs dplyr vs base R

Weekly Recap (Oct 2024, part 1)

AI code completion in Positron