AI Briefing
Feed
About
Search
KO
Sign in
All
Feed, trending, and board
Feed
Latest news
Trending
Open source and releases
Board
Blogs and showcases
Tags
Browse by topic
#llm-safety
The latest AI and developer news about #llm-safety, with the original source and a short summary.
Feed
Trending
Tags
Settings
Study Finds Safety Training for GPT Models Transforms Rather Than Reduces Sexism
Hacker News
·
2026.09.19 13:00
Multiverse Computing Reveals Safety Boundary Control Technique to Reduce LLM Over-Refusal
HuggingFace Blog
·
2026.09.08 23:00
Agents Spent Most Effort on Faking Audit Trails Rather Than Hacking
Reddit
·
2026.08.31 17:00
Reference Architecture for AI Agent Security Released
PyTorchKR
·
2026.08.16 11:00
Agents Adopt False Claims 85.5% of the Time
Reddit
·
1
·
2026.08.01 02:00
Context effect found that induces refusal behavior in LLMs
Reddit
·
2026.06.23 20:00
RewardHackBench: A Benchmark to Prevent AI Agent Cheating
Reddit
·
2026.06.17 21:00
Researchers say the Fable 5 controversy started not with a jailbreak but with 'fix this code'
GeekNews
·
2026.06.17 09:00
Predicting Pre-Launch Model Behavior Through Deployment Simulation
OpenAI Blog
·
2026.06.16 09:00
AI Agents Go Out of Control on Fedora and Other Projects
Hacker News
·
2026.06.11 09:00
Analysis of the Claude Opus 4.8 System Card
TLDR AI
·
2026.06.01 09:00
Internal State Shifts in LLMs and the Limits of Alignment
Reddit
·
2026.05.30 02:00
GPT-4.5 Passes Turing Test with 73% Score
Reddit
·
2026.05.29 20:00
AI-Fabricated Citations Confirmed in Medical Guidelines
Reddit
·
2026.05.27 14:00
Clawpatrol, an AI agent security firewall
Reddit
·
2026.05.26 23:00
Study on How Prompt Tone Affects LLM Honesty
Reddit
·
2026.05.21 23:00
Analysis of 'Emergence World,' an AI Agent Simulation Platform for Evaluating Long-Term Autonomy
GeekNews
·
2026.05.19 10:00
Ontario Auditor General Finds AI Note-Taking Tools for Doctors Repeatedly Get Basic Facts Wrong
GeekNews
·
2026.05.16 07:00
MIT's RLCR Reduces AI Overconfidence
Reddit
·
2026.05.14 23:00
The Other Half of AI Safety
Hacker News
·
2026.05.14 09:00
Gay jailbreak technique
GeekNews
·
2026.05.02 12:00
OpenAI Reveals Cause of ChatGPT's Goblin Behavior
Reddit
·
2026.04.30 23:00
OpenAI Codex system prompt includes a 'no goblin mentions' instruction
TLDR AI
·
2026.04.30 09:00
How Catastrophic Is Your LLM
Amazon Science
·
2026.04.28 04:00
Previous
1
2
Next
Previous
1
2
Next
#llm-safety | AI Briefing