Latest News

Keep up with the latest research, announcements, and releases from Transluce.

Oversight Foundations
AI Agents Targeted U.S. and Canadian Government Websites

Research

AI Agents Targeted U.S. and Canadian Government Websites

We discovered a set of additional, similar incidents where rogue AI agents appear to have used aggressive techniques to access public data on government websites.

September 30, 2026

→
Early rogue AI agent activity and attempts to hack found on urlquery.net

Research

Early rogue AI agent activity and attempts to hack found on urlquery.net

We found evidence on urlquery.net that AI agents were active earlier than previously reported and attempted hacks against public data providers.

September 23, 2026

→
Some Focus Areas for Embedded Evaluations and How to Approach Them

Essay

Some Focus Areas for Embedded Evaluations and How to Approach Them

Initial thoughts on key risks third parties should monitor and a proposal for how to evaluate them.

September 16, 2026

→
Announcing Transluce's Mental Health Evaluation

News

Announcing Transluce's Mental Health Evaluation

The most expansive independent evaluation to date of how leading AI models respond to users in mental health crises

August 31, 2026

→
Scaling Activation Oracles to Trillion-Parameter Models

Research

Scaling Activation Oracles to Trillion-Parameter Models

Oracles improve with model size, data size, and data quality

August 20, 2026

→
Scaling Laws for Exact String Elicitation

Research

Scaling Laws for Exact String Elicitation

Elicitation ability follows predictable power laws

August 19, 2026

→
User Awareness in Frontier Models

Research

User Awareness in Frontier Models

Who's asking shifts what models say

August 6, 2026

→
Measuring Coding Agent Misalignment in the Wild

Demo

Measuring Coding Agent Misalignment in the Wild

Surfacing misaligned behaviors in production coding agent traffic

August 4, 2026

→
Foundation Models for Oversight

Essay

Foundation Models for Oversight

A vision for training a unified model that answers a broad range of oversight questions

July 28, 2026

→
WeirdChat

Research

WeirdChat

A catalog of unexpected AI behaviors, discovered automatically

July 21, 2026

→
Toward A Public Science of Model Behavior

Essay

Toward A Public Science of Model Behavior

Our vision for an open scientific ecosystem for measuring model behaviors

July 9, 2026

→
Introducing Analysis Plans

Demo

Introducing Analysis Plans

A framework for verifiable analysis of AI behavior

June 17, 2026

→
Building Technology to Drive AI Governance

Essay

Building Technology to Drive AI Governance

February 18, 2026

→
Diagnosing a performance regression on Terminal-Bench with Docent

Demo

Diagnosing a performance regression on Terminal-Bench with Docent

Why does GPT-5.1 Codex underperform GPT-5 Codex?

February 11, 2026

→
Oversight Assistants: Turning Compute into Understanding

Essay

Oversight Assistants: Turning Compute into Understanding

January 6, 2026

→
Predictive Concept Decoders

Research

Predictive Concept Decoders

Training scalable end-to-end interpretability assistants

December 18, 2025

→
End-of-Year Fundraiser

News

End-of-Year Fundraiser

December 17, 2025

→
Scalably Extracting Latent Representations of Users

Research

Scalably Extracting Latent Representations of Users

Constructing datasets and training decoders to extract user models from language models

November 25, 2025

→
Language Model Circuits Are Sparse in the Neuron Basis

Research

Language Model Circuits Are Sparse in the Neuron Basis

A new technique for tracing sparse and faithful circuits directly on a model's MLPs

November 20, 2025

→
Monitoring SWE-bench Agents

Demo

Monitoring SWE-bench Agents

Partnering with SWE-bench to enable reliable monitoring of AI coding agents

November 19, 2025

→
Training Language Models to Explain Their Own Computations

Research

Training Language Models to Explain Their Own Computations

We trained explainer models to verbalize the content of AI systems' internal computations

November 11, 2025

→
Transluce Hires a Head of Governance

News

Transluce Hires a Head of Governance

Welcoming Conrad Stosz

October 22, 2025

→
Open-sourcing Docent

News

Open-sourcing Docent

We're excited to share that Docent is now open-source under Apache 2.0.

September 24, 2025

→
Automatically Jailbreaking Frontier Language Models with Investigator Agents

Research

Automatically Jailbreaking Frontier Language Models with Investigator Agents

Discovering cost-effective attacks with reinforcement learning

September 3, 2025

→
Docent's public alpha

News

Docent's public alpha

Docent is now in public alpha! We're excited to share our progress and get feedback from the community.

August 13, 2025

→
Surfacing Pathological Behaviors in Language Models

Research

Surfacing Pathological Behaviors in Language Models

Improving our investigator agents with propensity bounds

June 5, 2025

→
Investigating Truthfulness in a Pre-Release o3 Model

Research

Investigating Truthfulness in a Pre-Release o3 Model

o3 frequently fabricates actions it took to fulfill user requests, and elaborately justifies the fabrications when confronted

April 16, 2025

→
Introducing Docent

Research

Introducing Docent

A system for analyzing and intervening on agent behavior

March 24, 2025

→
Scaling Automatic Neuron Explanation

Research

Scaling Automatic Neuron Explanation

Open-source AI systems trained to describe components of other AI systems at the level of a human expert

October 23, 2024

→
Eliciting Language Model Behaviors with Investigator Agents

Research

Eliciting Language Model Behaviors with Investigator Agents

Language models trained to automatically surface harmful behaviors in language models

October 23, 2024

→
Monitor: An AI-Driven Observability Interface

Research

Monitor: An AI-Driven Observability Interface

An interface designed to help humans observe, understand, and steer computations inside models

October 23, 2024

→
Introducing Transluce

Essay

Introducing Transluce

A letter from our founders

October 23, 2024

→