Latest News

Keep up with the latest research, announcements, and releases from Transluce.

WeirdChat

Research

WeirdChat

A catalog of unexpected AI behaviors, discovered automatically

July 21, 2026

Toward A Public Science of Model Behavior

Essay

Toward A Public Science of Model Behavior

Our vision for an open scientific ecosystem for measuring model behaviors

July 9, 2026

Introducing Analysis Plans

Demo

Introducing Analysis Plans

A framework for verifiable analysis of AI behavior

June 17, 2026

Diagnosing a performance regression on Terminal-Bench with Docent

Demo

Diagnosing a performance regression on Terminal-Bench with Docent

Why does GPT-5.1 Codex underperform GPT-5 Codex?

February 11, 2026

Predictive Concept Decoders

Research

Predictive Concept Decoders

Training scalable end-to-end interpretability assistants

December 18, 2025

End-of-Year Fundraiser

News

End-of-Year Fundraiser

December 17, 2025

Scalably Extracting Latent Representations of Users

Research

Scalably Extracting Latent Representations of Users

Constructing datasets and training decoders to extract user models from language models

November 25, 2025

Language Model Circuits Are Sparse in the Neuron Basis

Research

Language Model Circuits Are Sparse in the Neuron Basis

A new technique for tracing sparse and faithful circuits directly on a model's MLPs

November 20, 2025

Monitoring SWE-bench Agents

Demo

Monitoring SWE-bench Agents

Partnering with SWE-bench to enable reliable monitoring of AI coding agents

November 19, 2025

Training Language Models to Explain Their Own Computations

Research

Training Language Models to Explain Their Own Computations

We trained explainer models to verbalize the content of AI systems' internal computations

November 11, 2025

Transluce Hires a Head of Governance

News

Transluce Hires a Head of Governance

Welcoming Conrad Stosz

October 22, 2025

Open-sourcing Docent

News

Open-sourcing Docent

We're excited to share that Docent is now open-source under Apache 2.0.

September 24, 2025

Automatically Jailbreaking Frontier Language Models with Investigator Agents

Research

Automatically Jailbreaking Frontier Language Models with Investigator Agents

Discovering cost-effective attacks with reinforcement learning

September 3, 2025

Docent's public alpha

News

Docent's public alpha

Docent is now in public alpha! We're excited to share our progress and get feedback from the community.

August 13, 2025

Surfacing Pathological Behaviors in Language Models

Research

Surfacing Pathological Behaviors in Language Models

Improving our investigator agents with propensity bounds

June 5, 2025

Investigating Truthfulness in a Pre-Release o3 Model

Research

Investigating Truthfulness in a Pre-Release o3 Model

o3 frequently fabricates actions it took to fulfill user requests, and elaborately justifies the fabrications when confronted

April 16, 2025

Introducing Docent

Research

Introducing Docent

A system for analyzing and intervening on agent behavior

March 24, 2025

Scaling Automatic Neuron Explanation

Research

Scaling Automatic Neuron Explanation

Open-source AI systems trained to describe components of other AI systems at the level of a human expert

October 23, 2024

Eliciting Language Model Behaviors with Investigator Agents

Research

Eliciting Language Model Behaviors with Investigator Agents

Language models trained to automatically surface harmful behaviors in language models

October 23, 2024

Monitor: An AI-Driven Observability Interface

Research

Monitor: An AI-Driven Observability Interface

An interface designed to help humans observe, understand, and steer computations inside models

October 23, 2024

Introducing Transluce

Essay

Introducing Transluce

A letter from our founders

October 23, 2024