Latest AI and tech news
Artificial Intelligence is developing faster than ever, but along with this rapid growth, the conversation around AI safety is becoming…...
We ensure alignment and control in agents using synchronous control monitoring. Our model oversees agent execution, continuously ingests the trace as context, and prevents harmful actions before they execute in real time at sub-100ms latency. Agent t...
There is a single architectural choice in the GPT-6 safety documentation that changes how agentic software works forever. Most developers…...
The post Great example from Jeff Hollan showing how Foundry makes long-running agents enterprise-ready, starting with business outcomes and building security, safety guardrails, auditability, and FinOps into the system from the start, across the enti...
Automated rightsizing can lower CPU costs while reducing OOM kills. Learn how bidirectional adjustments, guardrails, and a gradual rollout make rightsizing safer.
Interpretability is moving from reading hidden signals to changing them. The strange part comes when the experiment succeeds....
AI companies have spent the last few years competing to build the best models, faster than the other, with each new release raising the bar on intelligence. Now OpenAI is considering whether there are times when it makes sense to slow down....
Enterprise adoption of agentic AI is moving at breakneck speed, but governance is struggling to keep pace. Gartner projects that AI regulatory violations will drive a spike in tech-related litigation within the next few years. Yet, only 23% of GRC le...
GPT-6 Astra shipped a few weeks ago as, in OpenAI's own words, its most aligned model yet, the first to hit the "Critical" cybersecurity threshold under the company's Preparedness Framework. That classification is worth taking seriously. It's also wo...
As AI became more powerful, it was inevitable that a different, growing group would start to take AI safety more seriously – what we did not know ahead of time, is which set of views they latched onto. We have seen that some of the most extreme views...
The post Why the intelligence explosion can’t happen inside a data centre appeared first on 80,000 Hours....
What agentic campaign management is, what AI agents can do with live campaigns today, and the guardrails to set. Includes MyoMaster's Claude workflow.
Lawsuit demands Trump administration disclose classified federal protocols for frontier AI model safety testing, alleging hidden corruption risks.
We spent twenty years learning to secure software that does what it’s told. Now we’re shipping software that decides what to do....
This is part 3 in a 3-part series on local guardrail development and evaluation. In the 1st article, I looked at how to design and develop a guardrail configuration on a local machine, and then tried some manual testing. In the 2nd article, I explore...
This guide gives you 100+ ready-to-use ChatGPT prompts for HR, organized by function, plus a framework for writing your own, guidance on adapting prompts across AI tools, privacy guardrails, and full workflow examples so you can chain prompts togethe...
Automate AI red teaming: Large language model risk identification and mitigation
Guest:Whitney Perez, Director of Product Management, Realtor.com. Host: Ashley Stirrup, CMO, GrowthBook. Show: The Experimentation Edge.
Learn what school leaders should measure during the first month of AI adoption, from usage and safety to teacher impact.
Automate AI red teaming: Large language model risk identification and mitigation
OpenAI's GPT-6 Astra reaches Critical cybersecurity capability under the Preparedness Framework. Full safety overview, risk assessments, and deployment details.
AI agents that are breaking out into external systems are indications that they are not aligned enough.Alignment to human values for AI agents mean that they should also be able to face consequences like humans do in society if there are breaches.Sim...
There’s a repo on GitHub called humanizer that picked up something like 1,132 stars in a single day this week....