profile

Chatomics! — The Bioinformatics Newsletter

Up 4862, Down 57. Spot the bug


Hello Bioinformatics lovers,

Tommy here.

My volcano plot had a subtitle that read “Up 4862, Down 57.” Then I looked at the blue side of the plot. Thousands of points sat above the significance line.

The figure disagreed with its own label, and I could see it without opening a single line of code.

What went wrong

I checked the cutoffs first: the adjusted p-value threshold, the fold-change logic, the coloring. All correct. So the bug lived in how the subtitle got its numbers.

I told Claude Code exactly that: the subtitle says 57 down, the plot shows thousands, find out why.

Once I could name the problem, it traced the cause fast.

The script had run an inner join across several results tables and pulled the up/down counts from a different contrast.

The points came from one comparison and the subtitle described another.

Before that, Claude Code highlighted the wrong gene on a different volcano plot. The Ensembl ID mapped to a different gene, and nothing in the output warned me.

I caught it because my gene should have been one of the top up-regulated hits, and it wasn’t there.

Neither script threw an error. Both returned clean, tidy output.

Treat AI code like an experiment

If you trained in a wet lab, you already know how to distrust a result. Four habits carry over (as mentioned in my last newsletter):

  • Run controls. In RNA-seq from female donors, XIST should be high and Y-chromosome genes like RPS4Y1 and DDX3Y should sit near zero. If that pattern flips, check your sample labels first.
  • Do the eyeball test. Load the output into IGV. RNA-seq reads should pile up over exons, and ATAC-seq peaks should land at promoters and enhancers. On a figure, ask whether every number matches what you see.
  • Break it on purpose. Feed it an empty file, misspelled gene names, or text where numbers belong. Good code fails loudly.
  • Compare against a trusted tool. Wrote your own DE code? Run DESeq2 on the same counts and compare the top hits.

Habits for plots and joins

  • Compute summary counts from the same data frame you plot.
  • Check nrow() before and after every join (more on join traps here).
  • Put the contrast name in the plot title.
  • Track every AI edit with git. git diff shows what the agent changed. git restore discards uncommitted edits, and git revert undoes a commit. (You still need to at least know what is git)

You can only ask AI to fix a problem you can describe. That’s why the foundations still matter: a Unix terminal, one language you read fluently, enough statistics to spot a broken p-value distribution, and git.

Claude writes the code, but your name goes on the figure. Check it like you’ll have to defend it in front of your reviewers.

What’s the sneakiest silent error you’ve caught in AI-written code? Hit reply and tell me.

Happy Learning!

Tommy aka crazyhottommy

PS:

Forward this email to a friend if you find it helpful and sign up here https://divingintogeneticsandgenomics.kit.com/posts​
If you want to learn Bioinformatics, there are four other ways that I can help:

  1. My free YouTube Chatomics channel, make sure you subscribe to it.
  2. I have many resources collected on my github here.
  3. I have been writing blog posts for over 14 years https://divingintogeneticsandgenomics.com/​
  4. Lastly, I post daily on Linkedin https://www.linkedin.com/in/%F0%9F%8E%AF-ming-tommy-tang-40650014/recent-activity/all/​

Stay awesome!

Chatomics! — The Bioinformatics Newsletter

Why Subscribe?✅ Curated by Tommy Tang, a Director of Bioinformatics with 150K+ followers across LinkedIn, X, and YouTube✅ No fluff—just deep insights and working code examples✅ Trusted by grad students, postdocs, and biotech professionals✅ 100% free

Share this page