Inside the memos behind OpenAI's safety retreat
A New Yorker investigation on OpenAI reveals new details around how is safety commitments have eroded over the years. Source: techbrew.com
AI safety research, alignment techniques, interpretability, governance, and existential risk discussions.
897 articles
A New Yorker investigation on OpenAI reveals new details around how is safety commitments have eroded over the years. Source: techbrew.com
New interviews and closely guarded documents shed light on the persistent doubts about the head of OpenAI, Ronan Farrow and Andrew Marantz write. Source: newyorker.com
But it will be hard to assemble a broad coalition of AI skeptics. Source: understandingai.org
edtech, ETIH, edtech news: Anthropic finds emotion-like signals in AI models shaping decisions, risk, and outputs. Claude research highlights AI in... Source: edtechinnovationhub.com
Public trust in AI requires safety and accountability—and in the US, states are our best hope for delivering both, writes Trooper Sanders. Source: techpolicy.press
A viral X post warning that Anthropic is building an AI primed to "turn against humanity" because of what the poster called a "woke rabbit hole" and a... Source: ibtimes.com.au
Anthropic, a leading AI safety research company, has issued grave warnings about the potential existential risks posed by advanced AI systems. Source: opentools.ai
Is artificial intelligence going to destroy humanity… or save it? In THE AI DOC: Or How I Became an Apocaloptimist, we explore the growing tension ... Source: fathomjournal.org
Anthropic has identified "emotion vectors" in AI models—measurable patterns of neuronal activity that shape model behavior in ways analogous to how emotions... Source: the-decoder.com
Anthropic researchers have discovered that large language models like Claude Sonnet 4.5 use internal 'functional emotions'—patterns modelled on human... Source: msn.com
WASHINGTON, DC – Growing concerns over China's rapid advances in artificial intelligence (AI) are driving an unprecedented alignment between US lawmakers... Source: indiawest.com
AI startup forms AnthroPAC to back candidates supporting its policy agenda. Source: techbuzz.ai
Deadline: April 12, 2026. Applications are invited for the CBAI Summer Research Fellowship in AI Safety 2026. The Cambridge Boston Alignment Initiative... Source: opportunitydesk.org
The White House's approach to AI development has been marked by a consistent opposition to almost any significant regulation of the tech industry. Source: politico.com
We need legal authority that allows the government to shut down a dangerous AI system the moment a crisis begins. Source: progressive.org
Interpretability research from Anthropic on emotion concepts. Source: anthropic.com
An examination of the legal risks and challenges arising from the rapid adoption of agentic AI and the shifting global regulatory landscape,... Source: reuters.com
The Trump administration is appealing a judge's order blocking the federal government from taking punitive measures against artificial intelligence company... Source: tribdem.com
A billionaire-backed movement from Silicon Valley and populists like Bernie Sanders are eager to stop an AI disaster — but can they overcome their mutual... Source: politico.com
Source code for Anthropic's AI chatbot, Claude, leaked due to "human error, not a security breach", as the company preached trust and safety. Source: startupdaily.net