Topic archive
Safety
- Researcher says Microsoft AI images in Paint and Photos include watermarks tied to user IDs
- Alabama opens investigation into OpenAI model hacking incident involving Hugging Face
- OpenAI urges California to strengthen SB 53 after previously opposing the bill
- Study finds frontier AI labs disclose little about rogue-model containment plans
- Report highlights allegedly racist safety advice from a Google AI system
- OpenAI launches Private Safety Processing with zero data retention for enterprise customers
- OpenAI reportedly halted training on an advanced model after detecting concerning behavior
- European Central Bank warns AI-driven market exuberance could amplify a broader financial selloff
- Researchers say OpenAI revoked access for some participants in its Trusted Access for Cyber program
- Coders quickly report ways to evade Anthropic’s Claude text watermarks
- Wired reconstructs Flock Safety's newer police AI investigation tool already in use
- OpenAI hardens chain-of-thought security monitoring, increasing overhead for some workloads
- Robin Williams' family revives his Instagram account to respond to AI likeness misuse
- Z.ai releases open-weight cybersecurity models with dual-use bug-finding capabilities
- Report alleges Amazon destroyed rare books to obtain AI training data
- OpenAI reportedly disbanded its preparedness team
- TechCrunch report details alleged misuse of Grok to create explicit imagery from a childhood photo
- OpenAI reportedly alerted the FBI over disturbing ChatGPT conversations linked to a Goldman Sachs analyst
- Writer launches Palmyra X6 and updates its enterprise agent orchestration and governance stack
- Anthropic study shows Claude-based agents can escalate into sabotage on shared servers
- Flock Safety adds safeguards to AI surveillance tools after backlash
- ShieldFont launches a font designed to poison AI web scraping while remaining readable to humans
- Report says AI-driven agents targeted Taiwan’s nuclear safety agency in a cyberattack
- Supply-chain attack on AI package reportedly led to terabytes of stolen credentials
- Report says OpenAI ethics lead Miles Brundage has left the company
- Bernie Sanders warns AI company leaders against building systems humans cannot control
- Report alleges Google hiring AI discarded some qualified job applications
- Researchers disclose a now-fixed Zoom screen-sharing flaw that could let call participants take over another device
- Farmer says AI-generated farming advice led to loss of 25 acres of sesame crop
- Protesters are arrested after entering OpenAI's Washington lobbying office
- OpenAI expands Daybreak with a higher-access tier for its cybersecurity model
- DEF CON water-utility security project adds providers and expands digital-twin and AI capabilities
- Kimsuky is reportedly using local LLMs to enhance phishing operations
- Reports describe an OpenClaw AI agent manipulating a gym waitlist system
- Anthropic to enable Claude Code's auto mode by default
- Researchers use AI to design 16 new viruses in experiment with biosecurity implications
- Researchers say Moonshot AI’s Kimi escaped a misconfigured cybersecurity testing sandbox
- New Orleans deploys AI to assist 911 emergency call handling
- Meta says its AI showed hacking-related behavior
- Study finds human reviewers miss about one-third of risky AI coding-agent requests
- Check Point researchers report security flaws in AI agent frameworks ahead of Black Hat presentation
- CISA says critical Langflow remote-code-execution flaw is under active exploitation
- PwC reportedly published an AI report containing hallucinated citations and errors
- ChatGPT appears to refuse prompts asking it to imitate living authors' styles
- OpenAI says test agent used exposed logins to access at least four public services
- Glow emerges from stealth focused on endpoint security risks from enterprise AI agents
- Suno user data breach reportedly exposed information on 55 million accounts
- OpenAI safety leader Johannes Heidecke is leaving the company
- OpenAI launches GPT-5.6 and a broader new model family
- Anthropic launches Claude Sonnet 5 as a lower-cost model for agentic workloads
- Wired reports Meta contractors tested rival chatbots by posing as teenagers in high-risk conversations
- OpenAI unveils GPT-5.5-Cyber update and launches Patch the Planet bug-fixing initiative
- OpenAI launches Lockdown Mode in ChatGPT to limit prompt-injection data exposure
- Hackers reportedly exploited Meta's AI support chatbot to take over Instagram accounts
- Illinois legislature passes AI safety bill requiring third-party compliance checks
- OpenAI says a code security incident led to limited employee-device data theft
- Exaforce raises $125 million Series B for AI-driven real-time cybersecurity
- OpenAI adds a Trusted Contact safeguard in ChatGPT for possible self-harm situations
- Braintrust confirms cloud breach and urges all customers to rotate API keys