Science is safe: OpenAI new GPT-Rosalind fails 65% of actual lab tests
Turns out, memorizing Wikipedia isn't the same as doing actual science. OpenAI tried to prove its AI is ready for the lab, but their shiny new benchmark just ended up exposing how spectacularly these models choke when faced with real research.
The new benchmark, OpenAI's LifeSciBench, was built by 173 PhDs and biotech experts who wrote 750 practical tasks to see if AI can act like a real lab colleague. Instead of easy multiple-choice questions, the models had to analyze incomplete data, plan complex experiments, and evaluate research risks based on messy spreadsheets and charts.
When the results came in, the flagship GPT-Rosalind—specifically built for scientific work—completely bombed, failing 64% of the tasks with a strict pass rate of just 36.1%. Its older sibling, GPT-5.5, fared even worse, scraping together a sad 25.7% pass rate.
While OpenAI tried to spin the numbers by highlighting a softer "partial credit" score of 63%, the cold hard truth shows up when the AI has to deal with actual scientific artifacts. The moment a task requires analyzing an attached chart or data table instead of plain text, the pass rate for GPT-Rosalind plummets from 45.1% down to a miserable 28.1%.
The absolute worst performance happens when exact numbers are involved, where the model's accuracy bottoms out at 14.8%, and sequence structures don't fare much better at 24%. To ensure the test was fair, 453 independent PhD reviewers double-checked everything, agreeing that the tasks are indeed highly accurate reflections of daily lab struggles.
In a classic move of corporate cherry-picking, the comparison charts completely ignored Anthropic's Claude, which is known for its scientific precision. Instead, OpenAI preferred to grade its own homework, using GPT-5.5 as the automated judge to score all the rival models' answers.
The tech industry spent years promising that synthetic biologists would cure all diseases by next Tuesday, but the reality is that today's smartest models still get confused by a basic Excel spreadsheet. Watching a multi-billion-dollar neural network fail a high school level math check while trying to design a cancer drug is the perfect reality check for the AI hype machine.
Source: OpenAI
Comments
This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.