AI Agent Monitoring 🚀 Goodfire’s Inside-Out AI Safety
Discover Goodfire’s open-source AI agent monitoring: inspect model internals, boost AI safety monitors, and get security checks before production—learn trade-offs.

Can AI Monitors Catch Risky Agents for Less?
AI agents can spend hours taking actions, calling tools, and producing far more output than a person could reasonably review. That creates a real safety headache: if you want to know whether an agent is trying to bypass a restriction, exploit a system, or game its reward signal, someone (or something) has to watch what it does.
One common fix is to add a second AI model as a monitor. It reads the agent's activity and flags anything suspicious. That approach is straightforward, but it can get pricey - and slow - when the agent runs for a long time. Goodfire is pitching a different setup: monitor signals inside the model as it runs, and only bring in another AI system when those signals look worrying.
The company says this can make monitoring substantially cheaper. That's promising, especially for open models and hosted inference, but it doesn't automatically prove the underlying system is safe. Below is how the approach works, what Goodfire's reported results show (and don't), and what businesses should check before using it in production.
Why AI agents are difficult to monitor
A chatbot usually responds to a prompt and stops. An agent can operate through a longer chain: interpret a goal, make a plan, use tools, check results, then decide what to do next. It might write code, query a database, browse the web, or interact with another service. Each step is another chance for the agent to go wrong - or to take a path its operators didn't expect.
That changes the oversight problem. Reviewing one response is manageable. Reviewing every action in a multi-hour workflow - plus intermediate reasoning and tool interactions - is a different kind of operational burden. And a monitor that tries to read everything can end up adding its own cost and delay.
Why a second model can become expensive
The typical model-based monitor is pretty direct: send the agent's output or activity to another model, ask it to look for policy violations, then act on that judgment. It can work when the volume is limited and the risky behavior shows up clearly in the text.
At scale, though, the monitor may need to process a lot of material. That means more inference cost, more latency, and more testing for the monitoring model itself. If the agent generates thousands of steps, "add a reviewer model" can become an expensive default. And even then, some risks won't be obvious from what the agent prints.
There's a second problem: some failures show up before a final answer or action. If an agent is trending toward an unsafe move, catching it early could give operators more options than discovering it only after the action has already happened. That's the logic behind watching internal model signals instead of relying only on what the agent outputs.
Open models make deployment controls more important
Open models give developers more freedom to choose models, hosting, and modifications. That's useful - but it also means safeguards can vary. In a closed provider setup, some safety and monitoring may be built in by default. With open deployments, safeguards can be absent, reduced, or removed - intentionally or not.
The risk isn't limited to a hobbyist downloading a model. If a provider serves that model to many customers, an unsafe or compromised deployment can ripple outward. So for businesses, the question isn't just "did the model get safety training?" It's also "what controls exist where the model is actually running?"
Goodfire is aiming at that deployment layer. Its monitors are offered to Baseten customers, a platform that hosts and runs models for companies. It's one response to a broader issue: how to add oversight when customers can deploy and operate models outside a single lab's controlled environment.
How Goodfire's inside-out monitors work
Goodfire describes its system as a set of small detectors, or probes, that look at internal signals produced as a model processes each step. Instead of using a second model to re-read every output, the probes classify information the monitored model already generated during its computation.
A quick analogy: airport screening. The probes are like walk-through scanners that check people as they move through. If something trips a deeper check, a separate system can inspect further. In Goodfire's setup, that second look comes from another AI model.
Probes inspect internal activations
When a language model generates a token, it does calculations across its network. Those calculations create intermediate signals - often called activations. Goodfire trains probes to spot patterns in those activations that may correlate with risky behavior.
The core argument is that the underlying computation is already happening. The probe reads and classifies intermediate results rather than forcing a second model to process the whole interaction from scratch. In an interview on the MAD Podcast, Goodfire CEO Eric Ho described the cost advantage as reusing work from the model's forward pass and then running a classifier over internal activations.
That matters for deployment. This isn't just a cheaper "send a prompt to a smaller model" trick. The monitor depends on access to internal computations and on probes trained for specific patterns. It's also tied to how the model is hosted and instrumented, so it's not automatically a drop-in monitor for every model on every API.
A flagged signal triggers a closer review
A probe is meant to identify a selected signal - not to make every safety decision by itself. When it flags something, another AI model can inspect more closely. The customer can also decide what happens next, with options like:
- Logging the event for later analysis
- Sending it to a human reviewer
- Refusing the request or stopping the relevant action
The "right" response depends on the use case. A flagged event in a low-risk internal test might just be worth logging. Something that looks like an attempt to access sensitive data may warrant pausing and human review. A prohibited action could justify an automatic refusal.
This is an operational point worth keeping in view: detection and response are different design choices. A system that reliably identifies suspicious activity but can't pause, contain, or escalate it is basically an alerting system - not necessarily a real safety control.
Baseten customers can choose which risks to monitor. Goodfire lists categories like offensive hacking, misuse involving chemical or biological weapons, and reward hacking. Those categories aren't interchangeable, and strong performance in one area doesn't guarantee similar performance in another.
What the cost and test results tell us
Goodfire reports that its first monitor was built around Kimi K3, an open model. In the company's tests, monitoring about one million exchanges cost roughly $185 using its probes. Goodfire compared that with about $5,420 for a lower-cost model that checks every step and around $200,000 for a top-tier model doing the same kind of monitoring.
If those numbers hold for a relevant production workload, the cost gap could meaningfully change the economics of continuous monitoring. But these are company-reported test results, not a universal pricing chart. Treat them as a starting point to validate against your own model, traffic patterns, risk categories, and monitoring setup.
Cost is only one part of the comparison
The reported results suggest a possible cost advantage, but you still need a like-for-like comparison. Before turning these into procurement expectations, ask what the test counted as an "exchange," what level of review each approach performed, and whether the comparison includes hosting, integration, storage, and human review.
Also check what the monitor is actually evaluating. If it only flags a subset of patterns, it can't be compared cleanly to a reviewer model that assesses a wider range - unless the evaluation accounts for that difference.
Goodfire also says its probes caught 93% of malicious hacking sessions in its Kimi K3 tests. It sent 5.5% of harmless sessions for a second look. That second number matters because it shows review isn't restricted only to clearly harmful cases - some benign situations will still get escalated. That means extra work for reviewers and occasional interruptions to legitimate tasks, even if the rate looked low in their tests.
Finally, the company reports that running four probes at once added less than 2% to the time before the model started responding. That's a useful latency datapoint, not a guarantee. Model architecture, hosting choices, probe configuration, and traffic patterns could all shift the results.
Early detection could help during evaluation and training
Goodfire CTO and co-founder Dan Balsam has argued that internal monitors can identify risky behavior before it fully plays out, including during model evaluation or training. That could be helpful: catching a pattern during controlled tests is better than discovering it after a deployed agent already took an unwanted action.
Still, early warning isn't the same thing as reliable prediction in every setting. A probe is only useful to the extent that its signal generalizes to the model versions, prompts, tools, and environments where you deploy it. Change the model or its operating context, and the link between internal signals and external behavior might change too.
Goodfire isn't the only group looking at this. In January, Google DeepMind said its research informed the deployment of misuse-detection probes in Gemini. That suggests internal monitoring is a broader research and deployment direction, not a capability unique to one product. The key question for any implementation is how it's validated - and what operators can do when it raises an alert.
What businesses should check before deploying
An internal monitor can be a useful layer in a safety system, but it shouldn't be treated as a substitute for permissions, testing, logging, or accountability. The better question isn't "does the model have a monitor?" It's "what failures does this monitor actually help prevent, and what happens when it's wrong?"
Define the risk and the response first
Start with what an agent can do, not with the feature list. Identify the tools it can access, the data it can read, and what changes it can make. Then decide which behaviors should trigger a review, which should pause execution, and which must be blocked.
For example, an internal research agent that can draft code but can't execute it has a different risk profile than a customer-facing agent with access to production systems. You probably won't get away with one blanket threshold or response policy for both. Your monitoring setup should reflect what a false negative and a false positive would cost in each workflow.
A practical setup should answer:
- What is being monitored? Specify model versions, agent steps, tool calls, and risk categories.
- What does a flag mean? Decide whether it's "send for review" or "treat as a confirmed violation."
- What happens immediately? Choose whether it logs, pauses, escalates, or refuses.
- Who owns the alert? Assign a team and response time, rather than dumping alerts into an unmonitored queue.
- How is performance measured? Track missed incidents, unnecessary escalations, response time, and operational cost.
Validate against your own workload
A vendor's benchmark can help you decide whether to run a pilot, but it can't tell you whether the same probes will work on your model, your prompts, or your definition of harmful behavior. Use representative examples - safe edge cases and realistic adversarial attempts. Measure both what it catches and what it unnecessarily escalates.
If your workflow allows it, run in shadow mode. In shadow mode, the system records or classifies activity without automatically interrupting users. That gives the team a chance to compare alerts against human review before using the signals to block real production actions. For high-impact actions, shadow mode may not be enough as the only safeguard; access controls and approval gates still matter.
Also retest after meaningful changes: switching model versions, changing tool permissions, adding new agent capabilities, or updating system prompts. Past performance reflects the tested configuration - not a permanent safety guarantee.
Keep defense in depth
Internal probes address one slice of the problem: identifying suspicious behavior from internal model signals. They don't remove the need to constrain what the agent can actually do. A solid deployment can combine monitoring with least-privilege access, sandboxing, rate limits, tool-level validation, audit logs, and human approval for consequential actions.
Each control plays a different role. A sandbox limits how much damage an action can cause. A monitor can catch behavior that warrants investigation. A permission boundary keeps the agent from accessing what it doesn't need. Logging helps with incident review. If one layer fails, the others reduce the odds that a single mistake becomes a serious incident.
There are also questions about the monitoring system itself: how the probes are trained, how often they're evaluated, whether reviewers can understand alerts, and what happens when the system can't confidently classify a signal. The goal isn't a monitor that "always knows." It's a system where limits are visible and failures are contained.
Conclusion
Goodfire's inside-out monitoring targets a real weakness in the default approach of "use another model to review everything": as agents run longer and generate more steps, continuous oversight can get expensive fast. Using internal signals to spot suspicious activity - and then escalating selected cases - could make monitoring more practical for some hosted open-model deployments.
That said, the reported cost and detection numbers are worth testing, not treating like universal benchmarks. The probes should be evaluated against your models, tools, and risk scenarios, and their alerts need to connect to an operational response. This is one layer of control, not a replacement for constrained permissions or human accountability.
If you're evaluating AI agents for production, start by mapping what each agent can access and change. Then run a pilot with representative behavior, track missed and unnecessary alerts, and decide what the system should do when it flags risk. That's usually a more useful safety decision than picking a monitor based on cost alone.
For more detail on the launch and the company's claims, read the original TechCrunch report.