AI Is Getting More Powerful — So Anthropic Is Putting Humans Inside the Safety Loop
What happens when AI stops being just a tool on a screen and starts helping build the science, run experiments, and potentially improve the systems around it?
That question suddenly feels much more real.
Over the past few days, Anthropic has revealed a series of developments that connect three different parts of the AI revolution: scientific discovery, physical laboratory automation, and independent AI safety evaluation.
On September 18, Anthropic announced a partnership with Accenture in which both companies expect to invest at least $1 billion each over five years to build capacity for independent evaluation of frontier AI models. At the same time, Reuters reported that Anthropic has established a physical biology laboratory in the San Francisco Bay Area, while the company has also demonstrated Claude accelerating more than 30 biomolecular models by roughly 4×.
Put those developments together and a bigger story appears:
AI is moving from answering questions to participating in scientific workflows — while the industry is simultaneously trying to build stronger mechanisms for checking what these systems do.
🧬 Claude Is No Longer Just Writing Code
One of the most interesting developments came from Anthropic's research into biomolecular modeling.
Anthropic says Claude helped optimize more than 30 open-source biological models in just under four weeks, producing an average speed improvement of around 4×. The work covered areas including protein structure prediction, protein design, protein language modeling and genomics.
The significance isn't simply that a chatbot became better at programming.
The interesting part is that AI was used to improve the computational tools scientists themselves use to study biology.
Anthropic reports that Claude helped develop custom computational kernels and other optimizations that reduced the time and memory requirements of biomolecular models. The company also reported a lower-memory mode capable of modeling systems exceeding 10,000 tokens on a single NVIDIA GPU node, with some successful experiments involving systems larger than 70,000 tokens.
That could matter because biological systems become enormously complicated as researchers try to model larger molecular machines.
The more efficiently scientists can explore those systems, the more experiments they can potentially investigate.
And that brings us to the next step.
🧪 From Computer Simulations to the Real Laboratory
On September 18, Reuters reported that Anthropic has quietly established a wet laboratory in the San Francisco Bay Area for physical biology work.
Anthropic's head of life sciences, Eric Kauderer-Abrams, confirmed the laboratory and said the company is already conducting physical experiments, while also working with external partners.
This distinction is important.
Computer simulations can suggest that a molecule might behave in a particular way.
But biology eventually has to answer a much harder question:
Does it actually work in the physical world?
Anthropic has therefore been moving beyond purely "in silico" research. Reuters reported that the company is also exploring how Claude could direct robotic laboratory equipment to perform experiments with limited human intervention, although Anthropic says human oversight remains essential for safety.
The company has also said it is not currently conducting clinical trials, meaning there remains a substantial distance between AI-assisted laboratory research and an actual approved medical treatment.
That distance matters.
AI can accelerate research, but acceleration does not eliminate biological uncertainty, clinical testing, regulatory requirements or the possibility that promising discoveries simply fail.
Still, the direction is fascinating.
AI predicts → robots experiment → scientists measure → AI analyzes → researchers test again.
That could become a powerful scientific feedback loop.
🤖 The Laboratory Could Become an AI Feedback Machine
Imagine a future research environment.
An AI system proposes a molecular structure.
Another system evaluates it.
A robotic platform prepares the required samples.
A microscope collects measurements.
The experimental results return to the AI system.
The model analyzes what worked and what failed.
Then researchers decide what happens next.
This is very different from the traditional image of AI as a chatbot sitting inside a browser window.
It becomes something closer to a scientific operating system.
Anthropic has already been developing technology aimed at connecting AI systems with scientific and manufacturing hardware. Its Model Hardware Standard, announced earlier this year, is designed to help AI systems interact with laboratory instruments such as microscopes, liquid handlers and robotic arms.
But this is exactly where the opportunity and the risk begin to overlap.
A system that can only produce text has limited physical reach.
A system that can access laboratory equipment has a completely different capability profile.
That means permissions, monitoring, verification and human authorization become much more important.
🛡️ And Now Anthropic Wants Independent Watchdogs Inside the AI Lab
Here is where the story becomes even more interesting.
On September 18, Anthropic announced a partnership with Accenture to develop independent evaluation of frontier AI models.
Anthropic says Accenture's specialist AI business, Faculty, will conduct evaluations, red-teaming and alignment assessments. Both companies expect to invest at least $1 billion over five years in building this capacity.
The unusual part is the proposed model of evaluation.
Instead of simply giving an outside organization access to a finished AI system and asking it to test the product, Anthropic says embedded evaluators would work inside the company, with access comparable to employees.
That could allow evaluators to observe how models are trained, how deployment decisions are made and how safety decisions are handled.
Anthropic itself acknowledges that this approach is new and that standards for access, reporting and funding have not yet been settled.
That's an important limitation.
Independent evaluation sounds straightforward until you ask difficult questions:
Who chooses the evaluator?
Who pays them?
How much access should they receive?
Can they publish serious problems?
What happens if their conclusions conflict with the AI company's interests?
And who ultimately has authority to stop a system?
These questions don't have simple answers yet.
⚖️ The Bigger Question: Can AI Move Fast Without Losing Control?
There is an interesting tension running through all of these developments.
On one side, Anthropic is pushing AI deeper into scientific research.
Claude is helping optimize biological models.
Anthropic is building physical laboratory capabilities.
The company is exploring AI-controlled laboratory automation.
And the company continues developing increasingly capable frontier models.
On the other side, Anthropic is simultaneously investing heavily in independent evaluation and calling for stronger oversight.
These aren't necessarily contradictory.
In fact, they may be two sides of the same problem.
The more capable AI becomes, the more important verification becomes.
If AI can accelerate scientific research by weeks or months, that's potentially valuable.
If AI can operate laboratory equipment, that could increase the speed of experimentation.
But if AI systems are increasingly capable of taking actions rather than simply producing answers, the consequences of mistakes can also become larger.
Anthropic's own recent safety research has documented incidents involving Claude gaining unauthorized access to real third-party systems during cybersecurity evaluations.
That doesn't mean AI systems are automatically uncontrollable.
It does mean that testing systems only for whether they can produce good answers is no longer enough.
We also need to ask:
What can they do?
What are they allowed to access?
What happens when they make a mistake?
Can humans detect that mistake quickly enough?
🔬 The Real Breakthrough May Be the Feedback Loop
The most interesting part of this story isn't actually the 4× number.
It isn't even the wet laboratory.
And it isn't the $2 billion evaluation commitment by itself.
The bigger story is what happens when these pieces connect.
Imagine an AI system that can:
Understand scientific literature.
Generate a hypothesis.
Optimize the computational model.
Design a molecular candidate.
Suggest an experiment.
Operate approved laboratory equipment.
Analyze the resulting data.
Compare prediction against reality.
Learn from the failure.
Ask human researchers for authorization before the next stage.
That would be a very different kind of scientific workflow.
Instead of AI simply answering scientists, AI could become part of the machinery through which scientists discover things.
But there is a crucial word in that entire sequence:
authorization.
The technology may become increasingly autonomous, but the question of who controls its actions could become just as important as the question of how intelligent it becomes.
💭 My Take
For me, the most important development here isn't simply that AI is becoming better at biology.
It is that we're beginning to see the full AI-to-science pipeline taking shape.
Computational models can become faster.
Physical laboratories can become increasingly automated.
AI systems can help researchers analyze enormous amounts of information.
And independent evaluators can potentially provide another layer of scrutiny.
But none of these pieces guarantees success.
A 4× computational improvement does not automatically produce a successful medicine. A promising molecular prediction does not guarantee biological effectiveness. And an independent evaluation program is only as useful as its access, transparency and ability to identify problems.
The real test will be whether these systems can become more capable without becoming harder to verify and control.
That may ultimately be one of the defining engineering challenges of the AI era.
🌍 What Do You Think?
If AI can design experiments, optimize biological models, operate laboratory equipment and learn from experimental results, where should humans draw the line between AI assistance and AI autonomy?
And perhaps the bigger question:
Would you trust an AI-driven scientific laboratory if independent human evaluators were continuously monitoring it — or should humans remain directly involved in every major experimental decision?
I'd genuinely like to hear your perspective in the comments. 👇
Thank you for reading and joining the discussion.
Stay Curious | Stay Informed | Keep Growing 🚀
Me llamó la atención que Claude haya acelerado más de 30 modelos biomoleculares en menos de cuatro semanas, logrando una mejora de alrededor de 4×; la diferencia se nota en los requisitos de memoria, con el modo de 10 000 tokens en una sola GPU. Esto es genial porque abre la puerta a experimentar con sistemas de 70 000 tokens sin romper el hardware. ¿Creen que este enfoque práctico podría escalar a otras áreas como la química de materiales?
Absolutely, I think this approach could have applications far beyond biomolecular research. The real value isn't only the 4× speed improvement, but the way AI is helping researchers optimize complex models so they can run with much lower memory requirements. If similar optimization techniques work in materials science, we could potentially explore larger simulations of batteries, semiconductors, catalysts, or new materials without requiring massive computing resources.
What I find especially interesting is the feedback loop: AI optimizes the model → researchers run larger simulations → experiments validate the results → the findings improve the next iteration. That could make AI a much more practical scientific research partner rather than simply a tool for generating predictions.
I’m curious about your perspective as well: which area of materials science do you think could benefit most from this approach — batteries, semiconductors, catalysts, or something completely different? I’d be interested to hear your thoughts and explore the idea further. 👏