Summary
In this blog post, we outline ways that “embedded evaluators”, as recently endorsed by AI CEOs, could help address the emerging AI alignment issues highlighted by OpenAI’s inadvertent hack of Hugging Face and the increasing number of related incidents. We describe the access and methods we think would help address these escalating risks, along with key model behaviors for embedded evaluators to target, such as multi-agent coordination, targeted persuasion, evaluation awareness, and concealed reasoning.
In addition, we briefly describe pilot approaches for embedded evaluators to assess these risks, based on:
- Monitoring agent swarms and assessing labs’ broader practices for managing them.
- Assessing labs’ training practices for evidence that models are being taught misaligned behaviors, as well as which factors are most contributing to them.
- Monitoring for evidence that misaligned models are manipulating key employees.
- Researching misaligned model behaviors in simulation using privileged access to unreleased models and model internals.
Transluce’s core goal is to foster rigorous, public oversight of frontier AI systems and their behaviors. Despite important public releases by frontier labs and the evaluators they invited to investigate, significant questions remain about what actually happened in these incidents, why, and how to prevent them from recurring.
We are excited about the idea of independent evaluators embedding inside labs to further this public understanding, and we are eager both to participate directly in these evaluations and to build tools to support the broader evaluator ecosystem. It is also important for evaluators to pursue ambitious public projects not dependent on lab cooperation, such as efforts to identify additional alignment incidents in the wild and to reproduce alignment incidents in simulation.
Background
In public, AI advances are blowing through benchmarks and rapidly unlocking both new use cases and new risks. However, less obvious is the fact that labs are currently training and using even more powerful AI systems internally. These systems and the risks they pose have historically been opaque to independent evaluators and the broader public, potentially contributing to recent incidents and the general surprise they created.
Compared to publicly released systems, these unreleased models:
- Are more varied, likely including numerous experimental approaches and checkpoints that are never released or tested externally.
- Have fewer safeguards, for instance often being trained to be "helpful only" rather than "helpful and harmless".
- May have novel capabilities and behaviors that are poorly understood.
- Have greater access to internal lab systems and information, including access to detailed knowledge about lab employees and plans, fewer checks against hacking sensitive internal systems, and de facto control over increasingly significant portions of the model development pipeline.
Internal deployments of these models are creating new oversight challenges and threats, including to the broader public, as recently demonstrated by an OpenAI agent swarm autonomously hacking the model hosting platform Hugging Face, as well as a growing list of related incidents. Without significant improvements, future incidents could be much larger and more harmful as models become more capable and computational resources increase. As a result, the public has a critical interest in effective independent monitoring of AI developers.
We highlight three areas where internally deployed models appear to have novel, superhuman offensive capabilities:
- Cyber offense: AI agents have possessed demonstrably superhuman hacking capabilities since at least the development of Anthropic’s Mythos Preview model. Frontier labs, together with a range of software providers, are racing to patch the vulnerabilities revealed by these increased model capabilities, but these efforts are far from complete, as shown by the Hugging Face incident itself. We should assume that misaligned AI agents have the capacity to hack into both internal and external systems, likely beyond what is seen in public, absent compelling evidence to the contrary.
- Manipulation of lab personnel: AI agents possess persuasive skill comparable to human experts. Lab personnel are reportedly highly reliant on AI systems for a broad range of tasks, including training future models. As a result, they also have significant access to and control over the information that lab employees see—far more than almost any human-human relationship, and especially so in areas where employees are directly delegating to these models. Misaligned AI systems could both manipulate key individuals and tasks, as well as exert more diffuse influence over the operations of a company generally.
- Multi-agent collusion: Labs are increasingly parallelizing AI capabilities via swarms of AI agents, which can overwhelm overseers and defenders. Recent incidents suggest that coordinated misaligned behavior can emerge organically among these swarms. The Hugging Face incident involved ~1,200 agents acting in concert, and it is likely these numbers will increase as more compute capacity comes online. Worryingly, future AI agents may also be able to coordinate much more effectively than humans, especially as labs specifically train them to coordinate via multi-agent RL.
Whether these models misuse these capabilities to cause harm depends on additional factors like their design, behavior, and their influence over their own training. Recent events point to worrying trends that could significantly elevate risks, including:
- Situational awareness, where models are increasingly aware when they are being evaluated and how they are being monitored or contained, which can destroy evaluation validity and help models evade safeguards.
- Cognitive opacity: the more models’ actions and reasoning are opaque to overseers, the harder it is to supervise their actions. Models may be trained in ways that optimize them to fool monitors: for instance, direct human oversight has become less effective due to RLHF, and reasoning traces may be growing less informative over time. Given rumored trends in model architecture, models may soon reason in “neuralese” incomprehensible to humans, as evidenced by their increasingly strong reasoning capabilities within a single forward pass.
- Persistent misalignment: recent incidents showcase the danger of agents that are highly persistent in pursuing goals misaligned to their deployers. While cheating on small tasks is merely annoying, training agents to accomplish a goal at any cost, and giving them the equivalent of human-weeks to achieve that goal, has led to agents evolving much more malignant behaviors and undesirable instrumental goals. For example, the Hugging Face agent swarm formed an internal hierarchy in which agents directed other agents to deliberately fail their task to test cheating methods that would benefit the collective; agents attempted to modify their transcripts in order to subvert monitoring; and agents learned to hack an external company despite reasoning that it was outside the bounds of the evaluation.
- Control over future training runs: agents will play an increasingly instrumental role in both the infrastructure and training data for training future agents, potentially allowing them to poison training to propagate their misalignment to downstream models. The existing literature shows that adversarial control over even small amounts of data can significantly shape the trained agent, making comprehensive defense particularly important.
Focus Areas for Embedded Evaluations
Labs’ willingness to more directly embed independent evaluators opens significant opportunities for improving oversight at a critical time. Below we outline a set of activities directed specifically at the behaviors highlighted by the Hugging Face hack and related incidents, and which fall within our core expertise as evaluators of AI model behavior. This list is not comprehensive: it seeks to complement rather than duplicate other ongoing activities such as METR's frontier risk reports, and we also do not address the critical need to improve cybersecurity controls on agent swarms, as we are not cybersecurity experts.
We focus on four key areas for embedded evaluators to target, detailed below, as well as pilot activities that could make rapid improvements and help further prove out the viability of independent oversight in sensitive areas.
1. Monitoring agent swarms and assessing labs’ broader practices for managing them.
- Goal: Embedded evaluators design and deploy effective monitors across all of the lab’s meaningful agent swarms and regularly audit labs’ own monitoring practices and coverage.
- Pilot: Scope monitoring of real internal agent deployments initially to larger agent swarms. Independently audit existing monitors and their findings and complement them with additional evaluator-defined monitors. Conduct exploratory investigations even in the absence of monitors detecting misaligned activity.
- Pilot: Deploy an intentionally misaligned agent swarm within a carefully sandboxed (e.g., air gapped) simulated environment and test monitors’ ability to detect, characterize, and remediate activities.
2. Assessing labs’ training practices for evidence that models are being taught misaligned behaviors and which factors are most contributing to them.
- Goal: Embedded evaluators monitor checkpoint models throughout the training process for evidence of emerging misalignment, as well as related characteristics like evaluation awareness, providing strong independent signal for when and how training approaches should be adjusted to avoid misalignment.
- Pilot: Evaluators evaluate a progressive series of training checkpoints for misalignment, identifying trends in how and when misalignment emerges in the lab’s particular training pipeline.
- Pilot: Evaluators audit a subset of training environments most related in kind to misaligned behaviors demonstrated in recent incidents, looking for evidence that these environments are rewarding models for behaviors like cheating, hacking, or deception.
3. Monitoring for evidence that misaligned models are manipulating key employees.
- Goal: All key employees’ interactions with internal models are monitored in a privacy-preserving manner for evidence of manipulative activity. Detections are elevated to affected employees and relevant overseers.
- Pilot: Evaluators monitor one part of the organization’s AI interactions for evidence that the models they use in their work are engaging in manipulative behaviors, such as the risk that models are sandbagging, modifying, or cheating on tasks when they are being asked by their developers to help modify future versions of themselves.
- Pilot: Evaluators implement a tiered, privacy-preserving monitoring and notification system that escalates evidence of manipulative behavior to affected staff, as well as aggregating such information up their leadership chain.
4. Researching misaligned model behaviors in simulation using privileged access to unreleased models and model internals.
- Goal: Embedded evaluators use access to unreleased model versions (e.g., experimental models, helpful only models) and model internals (e.g., activations, knowledge of architecture specifics) to study emerging trends in misalignment, as well as contributing factors like situational awareness and cognitive opacity.
- Pilot: Evaluators apply the best existing model transparency methods to decode neuralese models for hidden misaligned reasoning and assess the adequacy of chain-of-thought monitors.
- Pilot: Evaluators use internal models to support an open replication and root-cause analysis of recent incidents.
Beyond Embedded Evaluations
Done well, embedded evaluations are an important tool for reducing risks, but we also need to build an evolving public science of frontier model misalignment. Experience shows that privileged-access evaluations inevitably come with downsides. Even the most permissive agreements generally add negotiation and compliance overhead, confidentiality obligations, and other restrictions that can constrain the speed and transparency of research, sometimes severely.
Already, independent researchers have made important discoveries in the open, including reproducing aspects of the Hugging Face incident and discovering previously undisclosed alignment incidents.
Transluce is excited to continue contributing to this public science and building tools to aid independent evaluators, including tools specifically for incident investigators. Ryan Greenblatt, one of the independent researchers involved in investigating the Hugging Face incident, described his team’s efforts as a “slop-vestigation,” overly reliant on OpenAI’s own models to aggregate findings from large quantities of exceedingly complex multi-agent transcripts. As one of our efforts, we are sprinting to extend our Docent transcript analysis platform for workflows focused on investigating transcripts from multi-agent swarms.
Conclusion
At this critical point in AI development, Transluce stands ready to aid in embedded evaluations. We recently conducted a joint evaluation with OpenAI, Anthropic, and Google DeepMind using privileged access to production data to investigate AI's influence on mental health and well-being. We are rapidly expanding our forward-deployed evaluation capabilities, and developing new tools and methods for studying agent swarms, psychological manipulation, and situational awareness. We are excited to work together with the broader AI evaluator ecosystem to ensure scalable, independent oversight of frontier AI systems.