OpenAI’s upcoming “Astra” model too powerful for current cyber safeguards

OpenAI says it has slowed some work on its upcoming model Astra after internal testing suggested it may have reached a new level of cybersecurity capability. The company says it “cannot rule out” that Astra meets its Critical threshold under its Preparedness Framework, its highest risk category for cyber capabilities. Under that framework, a model …

Read more

Sam Altman heads to Washington as AI security and China policy collide

OpenAI chief executive Sam Altman is set to meet senior Trump administration officials, senators and economists in Washington this week as policymakers debate AI security and access to Chinese models. Ashley Capoot and Kate Rooney report for CNBC that Altman plans to preview OpenAI’s upcoming model family and answer questions about cybersecurity, AI agents and …

Read more

Survey: AI security incidents rise as companies deepen integration

A new survey by device management company Jamf finds that companies using artificial intelligence extensively are significantly more likely to experience security incidents tied to that use. The survey polled 687 IT professionals and highlights a growing gap between how fast organizations adopt AI and how well they can govern it. Nearly three quarters of …

Read more

OpenAI introduces Lockdown Mode to shield users from prompt injection attacks

OpenAI is rolling out a new optional security feature called Lockdown Mode, designed to protect users from prompt injection attacks. Igor Bonifacic reports for Engadget that the feature is aimed at people and organisations handling sensitive data. Prompt injection is a form of social engineering targeting AI chatbots. Attackers hide malicious instructions on webpages or …

Read more

Analysis: How dangerous is Claude Mythos really?

Anthropic’s newest and most capable AI model, Claude Mythos, can autonomously find and exploit security vulnerabilities in virtually all major software systems. As I wrote previously, Mythos is currently only available to a select group of technology companies through Project Glasswing, Anthropic’s initiative to patch critical software before the capabilities become more widely known. Zvi …

Read more

Claude Mythos: Anthropic restricts its most capable AI model over cybersecurity risks

Anthropic has introduced a new AI model it considers too dangerous to release publicly. The model, called Claude Mythos Preview, can autonomously find and exploit security vulnerabilities in software. Instead of making it widely available, Anthropic is sharing access with a coalition of more than 40 organizations as part of an initiative called Project Glasswing. …

Read more

OpenAI removes “safely” from mission statement as it transitions to for-profit structure

OpenAI has removed the word “safely” from its mission statement, a change documented in its 2024 tax filing with the Internal Revenue Service. The deletion coincides with the company’s transformation from a nonprofit organization into a for-profit business. The original mission statement from 2022 and 2023 read: “to build general-purpose artificial intelligence (AI) that safely …

Read more

How Anthropic’s obsession with AI safety became its secret weapon against OpenAI

Anthropic has emerged as a formidable competitor in the artificial intelligence industry by focusing on enterprise customers and positioning itself as the most safety-conscious AI company. The approach appears to be paying off both commercially and in investor confidence, even as critics question whether the company can maintain its principles while racing to capture market …

Read more

Security flaws expose thousands of users on AI agent platforms

Two major security incidents have exposed vulnerabilities in the rapidly growing ecosystem around AI agents, revealing risks when artificial intelligence creates software without human oversight. Cybersecurity firm Wiz discovered a significant security flaw in Moltbook, a social network designed exclusively for AI agents, Raphael Satter reports for Reuters. The vulnerability exposed private messages between agents, …

Read more

×