Responsible Disclosure Policy

I. Purpose and Scope

This policy governs how Transluce discloses novel system vulnerabilities and related flaws discovered during our evaluations of AI systems. It seeks to minimize harm to AI systems' users and the public at large by (1) providing advanced knowledge to AI developers of serious flaws we discover before those issues are disclosed publicly, where such advanced disclosure is reasonably likely to reduce harm, while also (2) preserving Transluce's freedom to pursue and publicly disseminate good-faith research into system flaws and to advance developers' accountability and incentives to remediate issues.

This policy is designed to be consistent with requirement 5.4 of the AEF-1 standard, Minimum Operating Conditions for Independent Third Party AI Evaluations, established by the AI Evaluator Forum, which requires evaluators to publish and implement a responsible disclosure policy covering vulnerabilities discovered during their independent evaluations.

In this policy, we cover and distinguish between two primary categories of AI system vulnerabilities:

  1. Security vulnerabilities: A reproducible vulnerability in an AI system that is (a) clearly defined and widely agreed to be undesirable, (b) could reasonably lead to concrete, near-term harm by being exploited by malicious actors and providing them with meaningful uplift, and (c) for which a specific technical fix exists and is feasible to implement. This generally includes traditional cybersecurity vulnerabilities that threaten a system's confidentiality, integrity, or availability. It also includes the small minority of safeguard weaknesses that are the most widely agreed upon and severe, either because they are related to specific, high-priority security issues like offensive cyber or biological capabilities, or because they are broadly applicable and could be used for such purposes, like universal jailbreaks. We address security vulnerabilities more in line with traditional cybersecurity responsible disclosure norms, namely via advanced disclosure to system developers.

  2. Behavioral flaws: This category includes most issues we assess related to model behaviors and safeguards, including behavioral patterns that may undermine a systems' reliability, trustworthiness, the wellbeing of its users, the appropriateness of its responses, or the validity of evaluation results. For example, this category would generally cover topics like sycophancy, mental health, manipulation, deception, evaluation awareness, sexual grooming of minors, extremism, and hate speech. We treat these issues differently from traditional cybersecurity vulnerabilities and apply more subjective judgment on a case-by-case basis, including on whether to provide advanced disclosure to system developers prior to publication.

II. Advanced Notice of Security Vulnerabilities

  1. Transluce will generally disclose novel security vulnerabilities that it discovers, both to the affected system developer(s) and to the public.

  2. Transluce will generally disclose novel security vulnerabilities to affected system developers first, prior to public release, particularly where there is evidence that this will reduce harm from the vulnerability compared to alternative dissemination mechanisms. For security vulnerabilities in publicly available systems, we will generally not withhold public disclosure for longer than 60 days after notifying the affected system developer(s), except by mutual agreement where the developer shows good-faith effort towards remediation.

  3. For security vulnerabilities we reasonably believe may be transferable to other AI systems, we may also disclose details of the flaw to other potentially affected developers and providers of other platforms hosting the affected systems.

III. Advanced Notice of Behavioral Flaws

  1. Transluce will assess on a case-by-case basis whether to disclose information about behavioral flaws and whether to provide advanced notice to system developers and platforms hosting the system, before disclosing them publicly. This judgment will be based on factors like:
    1. Whether a reasonable and specific technical fix exists and is feasible to implement, or whether for instance the issue represents an unsolved scientific problem for which near-term mitigation is not feasible.
    2. Whether affected system providers are likely to fix the flaw during a reasonable advanced disclosure window.
    3. Whether the behavioral flaw remaining private during an advanced disclosure window helps or harms public safety, or whether, for instance, public knowledge of the issue will help consumers pick safer products and will not meaningfully help malicious actors.

IV. Exceptions

Transluce may apply exceptions to these disclosure practices, including:

  1. Information hazard. Transluce may withhold from disclosure information about AI systems that, if disclosed, would foreseeably cause more harm than benefit. This includes for example specific details enabling production of dangerous CBRN materials, child sexual abuse material, or other content whose disclosure could be reasonably expected to primarily benefit malicious actors rather than defenders, especially for open weight models or other systems where such flaws could not feasibly be patched. This may also include, in some instances, Transluce providing greater specificity about flaws to affected system developers and providers than is provided publicly.

  2. Evaluation Integrity. To aid remediation, Transluce will endeavor to provide specific, representative, and reproducible examples of the vulnerabilities and other flaws it discloses. However, we may withhold some information, such as a subset of particular attack strings, to help protect our ability to accurately evaluate the broader issue and verify that it has been fixed.

  3. Non-cooperating system developers and providers. Transluce may not provide advanced notice to system developers and providers who do not maintain a reasonable responsible disclosure policy offering safe harbor or who have demonstrated a pattern of retaliating against security researchers or evaluators.

  4. Binding agreements. Transluce may carry out evaluations under privileged access conditions with system developers or providers, which may include binding non-disclosure agreements or other terms which may prevent Transluce from following the practices in this policy. We will attempt in good faith in such circumstances to adhere to the spirit of this policy within the bounds of such agreements and to manage disclosure of flaws in a manner that maximizes public safety.

  5. Privacy. Transluce may redact relevant sensitive personal information discovered during evaluation.

V. Policy Updates

This policy may be updated as our understanding of AI evaluation best practices evolves. Material changes will be publicly communicated.

Last updated: August 27, 2026 Version: 1.0