“The model learns what you actually did, not what you meant to do.” — Nadina D. Lisbon
Hello Sip Savants! 👋🏾
The goal was to build clean training data for an AI model. To do this, four labs ran the same experiment. The results… four different answers. How could this happen when the catalyst and protocol were the same, something they had agreed to in advance? One of the biggest contribution was how hard each team shook or stirred the mixture[1][2]. This is what gets buried in the numbers feeding your own models.
3 Tech Bites
🔬One test, four results
SLAC brought together four labs. Its own, plus Penn State, Stanford, and UC Santa Barbara. The goal was to run the same test on the same rhodium-based catalyst, which would contribute to building clean data for an AI model. Each answer was different. A big part of it came down to how hard the mixture was shaken or stirred. The paper itself points to heat management as a key source of the variability [1][2]. Once the labs standardized further, the results lined up (Nature Catalysis, July 31) [2].
🧬 What clean inputs make possible
Clean inputs are the foundation of research. Stanford’s roundup points to Evo 2 and Biomni. Evo 2 is a DNA language model trained on 9 trillion base pairs. This is the largest model ever trained for biology. Biomni is a research agent that around 15,000 scientists have used to run 100,000 different workflows [3].
🤝 Where people still have the edge
It’s about wants versus possibilities. Research ideas produced by AI have been rated as more novel by researchers, but many ideas were considered impractical. Then, when the follow-up study took into consideration whether the research was possible, the ideas generated by the participants ranked higher.
5-Minute Strategy
🧠 Find the Stir speed in your Own Data
Spend five minutes searching for the identical gap in a number that is crucial to your system.
Pick one figure you and your AI tools both rely on, something like “active accounts” or “time to resolution.”
Open the two places it comes from, whether that’s two dashboards, two teams, or two systems.
Write, in your own words, the exact definition each one uses.
Mark where they part ways. Often it’s a different cutoff, or a timestamp someone set at a different moment.
Send that one-line mismatch to whoever owns the pipeline, framed as a question rather than a fix.
A model can’t tell a definition gap from a real signal. That is still ours.
1 Big Idea
💡 Made, Not Just Stored
The story of the four labs surprised me as my favorite science story of the summer. The labs performed the same experiment with the same catalyst. To ensure the data collected would be clean for their AI model, the labs agreed to follow the steps in a certain order. The labs ended up with four different results. Interestingly enough, the labs all attributed some of their different results to how much each lab shook or stirred the mixture. Variability in how a mixture is stirred or shaken is one of the most common variables in a reaction. The best part about this story is that there was actually no flaw in the science. The variability was due to stakeholder engagement.
Good news, right? The best these models can achieve isn’t actually determined by the models. The consistency of what we give to the models can be determined by people. The labs had to come to an agreement on the best stirring speed because of the variant in engagements. In my organization, the stirring speed is a gap that is usually defined as a difference in the recordings of the same event or in the way a manual step is performed in a different region.
Data is made, not just stored. Every number that feeds a model was produced by a person, a tool, and a habit, and the model faithfully learns all three. It can’t tell the difference between a real signal and a difference in how two teams work, because that judgment is ours. Which is a good argument for keeping people close to the data. The teams pulling ahead aren’t the ones with the most information. They’re the ones who care about how their information gets made.
When the inputs are clean, the payoff is real. The same field that produced the stirring study is using AI to read genomes, model human cells, and sort through telescope data no person could hold in their head [3]. Tools like Biomni, a biomedical research agent now used by roughly 15,000 scientists to run 100,000 different workflows, work alongside scientists rather than in place of them [3]. They take on the parts that never get tired, so people can spend their attention on the harder questions.
One more detail I keep thinking about. A pair of studies pitted AI against human experts on generating research ideas. The AI’s ideas were judged more novel. But when the follow-on study factored in whether those ideas could actually be carried out, the humans came out ahead [3]. That split seems about right to me. The machine widens what we can imagine, and we still supply the judgment about what will hold, which mostly comes down to agreeing on how hard to stir and writing it down.
If you’ve ever found a definition gap hiding inside a number everyone trusted, I’d love to hear how you caught it. I read every response.
P.S. If a peer is still bracing for cuts when the data points the other way, share this newsletter and help brew up stronger customer relationships.
P.P.S. If you found these AI insights valuable, a contribution to the Brew Pot helps keep the future of work brewing.
Resources
[1] For AI to drive science discoveries, highly reproducible data is key
[2] Bac et al., Nature Catalysis, 31 July 2026 (DOI: 10.1038/s41929-026-01559-y)
[3] From biology to astrophysics, AI expands the boundaries of discovery
Sip smarter, every Tuesday. (Refills are always free!)
Cheers,
Nadina
Host of TechSips with Nadina | Chief Strategy Architect ☕️🍵


