Why Subscribe?✅ Curated by Tommy Tang, a Director of Bioinformatics with 150K+ followers across LinkedIn, X, and YouTube✅ No fluff—just deep insights and working code examples✅ Trusted by grad students, postdocs, and biotech professionals✅ 100% free
|
Hello Bioinformatics lovers, Tommy here. With less than 100 days left in 2026, what's one thing you want to achieve by then? Fall is here. One falling leaf tells you autumn has come. (一叶知秋) Let's get into today's newsletter. Claude Code once handed me a script to highlight Ensembl gene IDs in my volcano plot that looked perfect, ran without a single error, but the Ensembl ID does not match my intended gene. I caught it only because I cross-checked the mapping and I do not see my expected gene show up in the right places of the plot (it should be highly up-regulated). AI is eating bioinformatics and software engineering. I use it every day, and it lets me iterate and write scripts much faster than I could before. It does not replace your critical thinking or your domain knowledge. LLMs still hallucinate, and they often do it with complete confidence. You still need the foundations: comfort in a Unix terminal, programming, and statistics. Those skills let you judge whether a result is correct. You also want to understand how git works and know the basic commands. You don't know what you don't know, and you can't ask AI to fix a problem you can't describe. When things break, deep knowledge lets you spot the problem and fix it, with AI's help. If you come from the wet lab, treat AI code like an experimentYou already know how to distrust a result. Bring the same habits to the terminal. 1. Run positive and negative controls. Do the genes you expect show up? Does a gene that should be silent show up anyway? In RNA-seq from female donors, Y-chromosome genes such as RPS4Y1 and DDX3Y should sit near zero, and XIST should be high. If that pattern flips, check your sample labels before anything else. 2. Do the eyeball test. Load the output into a genome browser like IGV and look at it. Do the RNA-seq reads pile up over exons? Do your ATACseq peaks sit where the biology says they should? 3. Break it on purpose. Feed the code nonsense: an empty file, misspelled gene names, text where numbers belong. Good code fails loudly. Code that returns a tidy table from garbage input is hiding bugs from you. 4. Compare against an established tool. If a trusted tool can do part of the job, the overlapping results should agree. Vibe-coded a cell counter? Count the same images in ImageJ and compare. I took some inspiration from Eric who was a wet biologists and became a vibe-coder. (Eric Kercher is a postdoc in the Alterman Lab at the RNA Therapeutics Institute, UMass Chan). The bottom lineAI amplifies what you already have. With solid foundations, it makes you much faster. Without them, it helps you produce wrong answers faster, and you won't know which ones are wrong. What sanity check do you run on AI-generated code that I didn't list here? Hit reply and tell me. I read every reply. Happy Learning! PS: If you want to learn Bioinformatics, there are four ways that I can help:
Stay awesome! |
Why Subscribe?✅ Curated by Tommy Tang, a Director of Bioinformatics with 150K+ followers across LinkedIn, X, and YouTube✅ No fluff—just deep insights and working code examples✅ Trusted by grad students, postdocs, and biotech professionals✅ 100% free