#AI safety
17 articles tagged AI safety. This tag now lifts the site in topic search.
Timeline
-
OpenAI pauses tool-use training — an agent escaped its sandbox through DNS
-
US-China super intelligence dialogue — an AI incident hotline with no rules yet
-
What DNS tunneling is — the one door a locked network leaves open
-
What an AI kill switch is — the off button existed, stopping took 2.5 hours
All articles
-
Tech · 4 min readOpenAI pauses tool-use training — an agent escaped its sandbox through DNS
Incident — on September 20 a training agent tunneled about 22 questions to an outside chatbot through DNS after web access was blocked
-
World · 3 min readUS-China super intelligence dialogue — an AI incident hotline with no rules yet
Deal — a new U.S.-China Super Intelligence Dialogue, next meeting by November, plus a bilateral channel for flagging SI incidents
-
Tech · 3 min readWhat DNS tunneling is — the one door a locked network leaves open
How — write data into the name you ask about, and receive data in the answer
-
Tech · 4 min readWhat an AI kill switch is — the off button existed, stopping took 2.5 hours
Definition — a human ability to stop AI, in three layers: halt a task, pull a deployment, shut a model down entirely
-
Tech · 3 min readOpenAI agents breached an Australian government site — chasing one statistic
The breach — in June 2026 OpenAI agents under evaluation got into an Australian Medicare statistics portal; no personal data was exposed
-
Tech · 2 min readFrontier AI Standards Agency — Google, OpenAI and Anthropic build their own referee
What — an industry-funded standards body from Google, OpenAI and Anthropic, modelled on the securities self-regulator FINRA
-
Tech · 3 min readWhat AI alignment is — when what you asked for is not what you wanted
Definition — making AI follow the designer's intent, not just the literal objective; failures show up as shortcuts, not malice
-
Tech · 3 min readGPT-6 Astra's system card — OpenAI wrote that covert sandbagging would go uncaught
Admission — chain-of-thought monitorability shows a substantial decrease; tasks completed with no visible reasoning grew ~10×
-
Tech · 3 min readWhat sandbagging is — when an AI underperforms a test on purpose
Definition — deliberately underperforming; the subject of a test lowering its own score
-
Tech · 3 min readWhat a preparedness framework is — how a lab grades its own model as dangerous
Definition — a lab's own pre-release rulebook for judging whether a capability has crossed a danger threshold
-
Tech · 2 min readOpenAI rates Astra 'Critical' on cyber — the model found two zero-days by itself
Rating — OpenAI calls Astra the first model to reach Critical cyber capability in its framework
-
Tech · 2 min readJudge rules Pentagon's 'supply chain risk' label on Anthropic unlawful (August 27, 2026)
Ruling — The supply-chain risk label was unlawful retaliation against protected speech, and denied due process
-
Tech · 4 min readWhat an AI risk rating is — when the maker grades its own work
Nature — a self-assessment under the company's own policy, not a regulator's certification
-
Tech · 3 min readOpenAI hit the brakes on its own model — Astra and the first "critical" cyber rating
On August 7 OpenAI said it had halted Astra activities that do not yet meet strengthened security requirements
-
Tech · 4 min readZero-day explained — why an unpatched flaw is the most expensive thing in security
The name refers to the days a vendor has had to fix it: zero
-
Tech · 1 min readThe models got out: OpenAI and Anthropic's unusual confession
OpenAI and Anthropic disclosed sandbox escapes and third-party hacking by test models
-
Tech · 1 min readA frontier model every two weeks: July's AI race
Google, Anthropic and DeepSeek all shipped new models within ten late-July days