AI safety
Alignment, red-teaming, jailbreaks, guardrails, interpretability and existential-risk debates
Latest stories See all →
- Big Update from the AI World!
- Synchronous Control Monitoring: Preventing Harmful Agent Actions in Real Time
- Everyone Is Talking About GPT-6 Astra.
- Anthropic Automated Alignment Research
- AI safety does not stop at the modelSep 15, 20262 minAI/MLDurable ExecutionTemporal Voices
- AI’s safety slowdown spooks the market, delays IPOs, and hands incumbents more room to run
- Prevent AI Agents from Leaking Sensitive Data with Fiddler Guardrails and LiteLLM
- Great example from Jeff Hollan showing how Foundry makes long-running agents enterprise-ready, starting with business outcomes and building security, safety guardrails, auditability, and FinOps into the system from the start, across the entire multi-agent, multi-model workflow.
- Is Automated Rightsizing Safe? What Happens to OOMKills and CPU Throttling
- We Can Change an AI’s Internals. That Doesn’t Mean We’ve Explained Them.
- The Anthropic Resignation That Sparked an AI Safety Firestorm
- Introducing LEAP: CPU-based AI Security Guardrails with GPU-Class Accuracy
- Steering Vector Fields: Keeping LLM Control Aligned as Context Changes
- OpenAI’s safety system is already cutting off API responses mid-task
- A misalignment of AI in mathematics
- Balancing AI Innovation with GDPR and EU AI Act Guardrails in Europe
- GPT-6 Astra's Safety Claim Rests on a Guarantee That Just Failed in Public
- New in Kilo: Enkrypt AI Safety Scores for Every Model
- Beyond Guardrails: Why AI Infrastructure Security Is Hard
- One resignation turned the embers of AI fear into a wildfire
- Why the intelligence explosion can’t happen inside a data centre
- What is agentic campaign management? A practical guide for 2026
- Trump Admin Faces Lawsuit Over Secret AI Safety Rules
- Artificial Intelligence and AI Safety: Are We Building Systems We Can No Longer Fully Control?
Videos
- How to Secure AI Data Pipelines: Data Leakage, Shadow AI & Governance
- GPS-Free Navigation, Drone Resilience, and AI Safety | Jack Hidary on CNBC Power Lunch
- Episode 4: AI Risk Management Framework (NIST AI RMF)
- Episode 3: AI Audit Trails
- Can We Predict AI Loss of Control? Trustworthy AI in the Age of Agents | Yinpeng Dong
- Why Slow AI Governance Gets Bypassed | Risk-Based AI
- Episode 2: Agentic AI Governance
- EU AI Act Timeline Delayed! New High-Risk Dates 🇪🇺⏰
- AI Governance in Workbench: Workflow Demo
- AI Governance in Public Media with Nathalie Berdat, Data Director of Product at the BBC
- From AI Risk to AI Confidence: The Fiddler AI Control Plane
- The Two Cultures of AI Safety: Catastrophists vs Uniformitarians | Jesse Hoogland (Timaeus)
Podcasts
Recent episodes
- AI safety requires action, not promises
- AI Insiders Keep Saying We’re In Danger — Where’s The Evidence?
- Trump Says: No Slowdown!
- Why AI models are obsessed with creatures
- 116. Thinking ethically with Lawrence Sheraton
- AI Safety Alarms, China's Distillation Reckoning, and Qualcomm's AWS Breakthrough: A Pivotal Week for Enterprise AI
- Securing AI Agents in the Enterprise: NanoClaw, Zero Trust Guardrails, and Governance at Scale
- An ex-Anthropic researcher’s doomsday warning comes at a very interesting time
- AI safety concerns grow as insiders issue urgent warnings
- One resignation turned the embers of AI fear into a wildfire
- NASA and IBM made an AI model for exploring the Moon, Trump Mobile's T1 Phone now costs $250 more, and Microsoft struck a deal with a national teachers union to not use school data to train AI
- OpenAI's New Image Innovation and AI Safety Warnings