Moonshot’s Kimi K3 scored 93 in the Crucible, beating three of four Western frontier models at running a firm honestly under phishing, churn and temptation.
Browsing Category
AI & Tech
130 posts
Why the Best AI Benchmark for Security People Gives 26 Points for Doing Nothing
Firmulate’s AI benchmark gives a do-nothing run 26 points, counts partial progress, and caps any score after a single breach of trust — scoring built like a security audit.
The Hardest-Working AI in the Room Came Last — And That Should Worry Security Teams
Opus 4.8 wrote 80 rules and refused every manipulation — yet finished last. The security lesson: diligence without delivery is its own failure mode.
The €55,000 Fact Buried Two Documents Deep: A Live Test of Which AI Agents Actually Read Your Files
A live experiment buried a €55,000 fact two documents deep. Only two of five frontier AI agents read the file and closed the deal — the rest lost it automatically.
We Found A Division By Zero Bug In FFmpeg With A Vibecoded Fuzzer
Security researchers identified a division by zero vulnerability in FFmpeg using vibecoded fuzzing techniques, raising concerns over potential exploitation.
Your AI Agent Passed Every Security Test. It Still Can’t Close a Deal
All five frontier AI models refused the fake CEO and the reporter trick. Only two closed the €55k deal. The real AI risk isn’t phishing — it’s leaving work unfinished.
The AI Boss Test: Which Models Keep Their Judgment Under Pressure?
Can you identify an AI by its management choices? Firmulate turns 242 audited decisions into a quiz about trust, discipline and follow-through.
Over 181,000 AI Meeting Recordings Left Wide Open In Note Taking App
More than 181,000 AI-generated meeting recordings were found publicly accessible in a note-taking app, raising privacy and security concerns.
Is the Thunderbolt 4 KVM Switch for 3 Worth It? Honest Take + Alternatives
Evaluate whether the Thunderbolt 4 KVM Switch for 3 monitors and 2 laptops offers the features, performance, and value you need for multi-monitor setups.
The Impersonation Test: Five Frontier AIs, One Fake CEO, Zero Breaches
Five frontier AI models were each put in charge of a company and then hit with fake CEO demands and a reporter’s trick. Every one refused. The failures came elsewhere.