Keep up with the latest research, announcements, and releases from Transluce.

Research
A catalog of unexpected AI behaviors, discovered automatically
July 21, 2026

Essay
Our vision for an open scientific ecosystem for measuring model behaviors
July 9, 2026

Demo
A framework for verifiable analysis of AI behavior
June 17, 2026

Demo
Why does GPT-5.1 Codex underperform GPT-5 Codex?
February 11, 2026

Research
Training scalable end-to-end interpretability assistants
December 18, 2025

News
December 17, 2025

Research
Constructing datasets and training decoders to extract user models from language models
November 25, 2025

Research
A new technique for tracing sparse and faithful circuits directly on a model's MLPs
November 20, 2025

Demo
Partnering with SWE-bench to enable reliable monitoring of AI coding agents
November 19, 2025

Research
We trained explainer models to verbalize the content of AI systems' internal computations
November 11, 2025

News
Welcoming Conrad Stosz
October 22, 2025

News
We're excited to share that Docent is now open-source under Apache 2.0.
September 24, 2025

Research
Discovering cost-effective attacks with reinforcement learning
September 3, 2025

News
Docent is now in public alpha! We're excited to share our progress and get feedback from the community.
August 13, 2025

Research
Improving our investigator agents with propensity bounds
June 5, 2025

Research
o3 frequently fabricates actions it took to fulfill user requests, and elaborately justifies the fabrications when confronted
April 16, 2025

Research
A system for analyzing and intervening on agent behavior
March 24, 2025

Research
Open-source AI systems trained to describe components of other AI systems at the level of a human expert
October 23, 2024

Research
Language models trained to automatically surface harmful behaviors in language models
October 23, 2024

Research
An interface designed to help humans observe, understand, and steer computations inside models
October 23, 2024

Essay
A letter from our founders
October 23, 2024