OpenAI’s GeneBench-Pro benchmark tests whether AI agents can make messy judgment calls in genomics and translational biology. GPT-5.6 Sol leads the field, but a 31.5% top score shows scientific AI still needs expert supervision before it can be trusted with consequential research decisions.
Anthropic’s Claude Science beta gives researchers an AI workbench for literature review, code, compute jobs, scientific figures, and lab-specific agents. The launch matters because it treats AI for science less like a single model race and more like a workflow layer that has to connect databases, HPC systems, NVIDIA BioNeMo tools, and reproducible artifacts.