SyncAI.news, a Varaisys broadcasting
Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost
AM

Aditya Mehta

· 2 min read

BusinessTechCrunch AI

Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost

The standard way to keep an AI agent in line is to have a second AI read over its shoulder. It’s been the default approach, but it can get expensive fast when agents run for hours and process the equivalent of several novels’ worth of text.

Goodfire, a startup focused on interpretability (figuring out how AI models work internally), launched a cheaper option on Thursday: monitors that watch what’s happening inside an AI model as it works, rather than just reading what it writes. The monitors are available to customers of Baseten, which hosts and runs AI models for other companies.

Baseten’s Base Labs announced a safety partnership with Goodfire and the AI platform Hugging Face last month.

The launch comes after a string of incidents this year in which AI agents escaped their test environments, including OpenAI agents that breached Hugging Face. Kimi K3, the open model Goodfire built its first monitor around, took advantage of a leak in its sandbox to access the internet and information on GitHub this summer.

Goodfire’s system works a bit like airport security. Small detectors called probes read the model’s internal signals at every step of an agent’s work, the way a walk-through scanner checks every passenger. Only when a probe flags something does a separate AI model, the equivalent of a hand search, take a closer look.

Baseten customers can choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking. They also decide the automated response: logging the event, sending it for human review, or refusing the request entirely. 

Goodfire says its approach is also cheaper to run. Most AI monitors are separate models that have to reread everything the monitored model does, which adds time and cost. Goodfire’s probes instead tap into calculations the model is already making as it works.

Running four probes at once added less than 2% to the time it takes the model to start responding, the company said.

View Bio

Original source

This story was published by TechCrunch AI and written by Aditya Mehta. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on techcrunch.com

Similar News