There are over 6,000 abliterated AI models sitting on Hugging Face right now. That number alone tells you why this announcement matters. On Wednesday, Baseten’s research group Base Labs announced a new safety infrastructure standard for open-weight models, partnering with Hugging Face and interpretability startup Goodfire AI to build evaluation and monitoring tools into the foundation of how models are built and served.
Abliteration is the technique at the center of this. It refers to the process of stripping safeguards out of open-weight models after the fact, a method that has become increasingly common and increasingly easy. When anyone can download a model and remove its guardrails, the safety work done during training can be undone in minutes. That’s the problem Base Labs is trying to get ahead of.
The core argument from Baseten is that openness is actually an asset here, not a liability. “Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source,” the company wrote on X. It’s a direct counter to the narrative that open-weight models are inherently less safe than proprietary ones from companies like Anthropic or OpenAI. Whether that argument holds up depends entirely on whether the safety infrastructure being built is any good.
That’s where Goodfire comes in. The startup specializes in model interpretability, essentially opening the black box to explain why models produce the outputs they do. Goodfire’s framing of the partnership gives the clearest signal of the division of labor: “Safety must be built into open models and provided by those who serve them.” Baseten is the inference provider doing the serving. Goodfire is the most likely candidate for the “built into” part. Technical specifics haven’t been disclosed yet, which is a gap worth watching.
Both companies have the capital to back serious research. Baseten raised a $1.5 billion Series F in June at a $13 billion valuation. Goodfire closed a $150 million Series B led by B Capital earlier this year. These are not underfunded side projects.
Base Labs is also putting out an open call for developers to contribute to the framework, which signals this is being positioned as a community standard rather than a proprietary product. That matters because adoption is the real test. A safety standard nobody uses is just a white paper. The broader open-source AI ecosystem, including Meta’s Llama lineage and the growing field of fine-tuned derivatives, would need to buy in for this to actually shift how open-weight safety gets handled at scale.



