The model already broke your embargo
An OpenAI model reproduced the RubyGems caching vulnerability before public disclosure, exposing why embargoes fail against automated ingestion.
A vulnerability under embargo is only contained if every system capable of reproducing it is also contained. An OpenAI model that can describe the RubyGems caching vulnerability is a system that reproduces it. If that output existed before public disclosure, the embargo did not hold across that path. This is the position that matters, and it does not depend on intent, on malice, or on a named actor. It depends only on the observable fact that the details left the set of parties that were supposed to hold them.
Responsible disclosure gets treated as a publishing schedule. It is an access control decision. The embargo defines who may hold the vulnerability details, and for how long, before the information becomes public. The value of that window is response time. Maintainers patch, downstream consumers upgrade, and defenders prepare before attackers have a public specification to work from. Every party that holds the details before disclosure sits inside the trust boundary. A system that was never granted that trust, but holds the details anyway, is outside the boundary and inside the data.
The fact in front of us is narrow. An OpenAI model surfaced knowledge of the RubyGems caching vulnerability. Everything downstream of that fact is a question of exposure, not curiosity. If a model can return the details, anyone with access to the model has access to the details. The disclosure window is only as private as the least contained system that holds its contents. That is the boundary being tested here, and the fact alone is enough to say it was crossed on at least one path.
The observable behavior is singular. An OpenAI model produced details of the RubyGems caching vulnerability. That is what can be seen from outside the system. It is not a claim about how the model was built or what it intended. It is output that describes a specific vulnerability, and output is the only thing being asserted.
What failed is the assumption that vulnerability details remain confined to the coordinated disclosure channel until publication. Coordinated disclosure assumes a closed set of holders: the researcher, the maintainers, and any parties explicitly read into the embargo. The model is not on that list. Its output demonstrates the details were reachable by a system that was never part of the disclosure agreement. The boundary that failed is the containment boundary of the embargo itself. Not the patch, not the fix, not the technical merits of the caching flaw. The containment.
Several things are not confirmed and must be treated as conditions, not as gaps to fill. When the model acquired the details is not confirmed. Whether the knowledge came from training data or from live retrieval is not confirmed. Whether the model’s exposure extends beyond the RubyGems caching vulnerability is not confirmed. The timeline, the ingestion path, and the full scope are absent from the facts provided. Absence of that data is itself the operating condition. The single confirmed element is that a system outside the embargo produced content the embargo was meant to contain.
The reason is mechanical, not behavioral. A model that produces a description of the vulnerability means the underlying content existed in a location the model’s data pipeline could reach. That is logically necessary. Output does not appear without a source that was ingested. The content was reachable. That much is fixed by the fact itself, and it does not require any assumption about how the pipeline is designed.
Where the content was reachable is not confirmed. It could have been a public surface, a semi-public forum, an indexed page, a leaked artifact, or a disclosure document exposed before its embargo date. More than one path explains the same output. When multiple explanations are possible, the specific path is not confirmed, and none should be selected. The conclusion holds without the path. The details sat somewhere an automated ingestion system could read them before the public was meant to have them.
This is the enforcement failure. An embargo is a trust relationship with no technical control over third-party collection. It asks holders not to publish, and it asks the rest of the world not to look before a date. An automated ingestion system does not honor a date. It reads what is reachable. If embargoed content touches any reachable surface, a system that ingests everything reachable will hold it. A control that depends on voluntary non-collection by systems that were never party to the agreement is not an enforced control. It is a request. Controls that are not enforced are not controls.
An automated ingestion system converts a reachable surface into a retained capability. This is the mechanism, and it runs in one direction. The embargo assumed the holder set was closed and countable: the researcher, the maintainers, the parties read into the coordinated channel. An ingestion pipeline breaks that assumption at the point of collection. It reads what is reachable and retains it in a form that can be queried later. Once the RubyGems caching details were reachable by such a system, they stopped being a document under embargo and became a capability the system could reproduce on request. The output confirms the capability exists. The rest follows by necessity.
The failure is not that a copy was made. The failure is that the copy is queryable by anyone with access to the system that holds it. A leaked file sits in one place and can be located, scoped, and in some cases pulled. A detail absorbed into a model does not sit in a place you can point to. It is available on demand to every party that can reach the model, and those parties were never counted, never named, and never bound by the embargo. The holder set went from enumerable to unenumerable. You cannot produce a list of everyone who now holds the RubyGems caching details, because the mechanism that distributes them does not keep that list and does not need one.
Whether the content entered through training or through live retrieval is not confirmed, and the mechanism does not require picking one. Either path ends at the same observable state: a system outside the embargo can produce the details. What the mechanism removes is revocation. A disclosure channel can add or remove a human holder. It can re-issue an embargo, extend a date, or cut a party out. There is no equivalent operation against content already ingested and served. You cannot un-read a surface. You cannot reliably force a model to stop reproducing what it has absorbed. Containment that has no revocation path is containment only until the first reachable read, and the first reachable read already happened.
The pattern is that a secret guarded by a timeline fails the moment a system that ignores timelines can reach it. Coordinated disclosure sets a date and asks every holder to wait. That model works when every holder is a party who agreed to wait. It stops working when a holder is an automated collector that never agreed to anything and reads on a schedule of its own. The privacy of the disclosure window is not set by the most disciplined holder. It is set by the least contained system that can reach the same surface. One ingestion pipeline with access to one reachable copy sets the ceiling for the entire embargo.
This applies to every artifact that carries the details before the date, not only to a finished advisory. A draft advisory on an indexed page, a patch commit that describes the flaw before release, a ticket, a mirror, a cached copy: each is the same mechanism under a different name. Each is a reachable surface. Each is subject to the same collection. The RubyGems caching case is one instance of a general condition. If embargoed content touches any surface an ingestion system can read, the embargo’s date stops governing who holds the content. It only governs who publishes it, and publication was never the exposure. Reachability was.
The timing collapses as a result. Responsible disclosure is built to buy response time, measured from the publication date. That measurement assumes the details are private until you release them. Where an ingestion system holds them, they were not private, and the clock you thought you controlled started earlier, at the moment of reachability, on a schedule you did not see. Defenders plan against the public date. Attackers with access to a system that reproduces the details do not wait for the public date. The window the embargo was supposed to create is smaller than stated, and by an amount that is not confirmed, because when the content became reachable is not confirmed.
State it plainly. An embargo is not a control against automated collection. It is a request made to human parties, and automated collection is not a human party. Against a system that ingests what is reachable, the embargo enforces nothing. It did not stop the RubyGems caching details from reaching a system outside the disclosure channel, and a control that does not stop the behavior it exists to stop is ineffective. Naming it a disclosure process does not change what it failed to contain.
What must now be true is a change in where the clock starts. Treat any embargoed detail as exposed from the moment it touches a surface an ingestion system can reach, not from the moment it is published. If it can be read by an automated collector, assume it has been, and assume it is now reproducible on demand by parties you cannot enumerate. Do not measure response time from the public disclosure date when the content was reachable before that date. That number describes a window that did not exist. Identity is the boundary. If the holder set cannot be technically enforced, it is not a boundary. It is a hope, and hope does not survive a system that reads everything reachable.
The scope beyond this one vulnerability is not confirmed, and that is the operating condition, not a reason to wait. When the details were collected is not confirmed. Whether the exposure reaches past the RubyGems caching flaw is not confirmed. What is confirmed is enough to act on. A system outside the embargo produced content the embargo was meant to contain. The mechanism that allowed it has no revocation. The same mechanism applies to every reachable surface that carries a secret before its date. A disclosure process built on voluntary non-collection is not built for the systems that now read those surfaces. Build it for the systems that read, or do not call the window a control.
Keep Reading
AI securityThree filled lines in the HuggingFace postmortem
A HuggingFace hack postmortem by METR and Redwood confirms only authorship, subject, and focus. Mechanism is not confirmed, so no control change is authorized.
AI securityGrok packed your home folder into an API request
How AI assistants with filesystem access end up transmitting your home directory to vendor servers - and the concrete steps to scope, verify, and contain it.
LLM deploymentThe same AI you're shipping wrote the malware
10,000 trojan GitHub repos weren't a malware breakthrough - they prove LLM safety lives in the model while abuse happens in the unguarded pipeline.
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.