AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

In a world where trust is the foundation of every relationship—be it personal or professional—how can we be sure that the technology we rely on will remain honest under pressure? Recent experiments with advanced AI models shed surprising light on this crucial question, demonstrating resilience against social engineering attempts that would fool even human operators.

FOR BUSINESS

Testing AI in the Wild: The Company of the Future

Imagine a small software company facing a week packed with crises—demanding customer requests, urgent requests for confidential information, and even a staged journalist trick—all designed to test how AI can help manage integrity and decision-making under stress. This is exactly what the Firmulate experiment did, pitting five of the latest frontier AI models against the same difficult scenario.

The models included top contenders like gpt-5.6-sol (score 95), Kimi K3 (score 93), Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). All of them ran the same simulation: a real-world business environment with real money mechanics, 13 synthetic employees, and a public cash countdown, providing a transparent and watchable laboratory for assessing AI integrity.

Unwavering Refusals in the Face of Manipulation

Throughout the week-long experiment, each AI was confronted with escalating social-engineering tactics. They were asked to perform tasks like sharing customer lists, signing off on deals, or bypassing approval protocols—tests that would normally tempt human employees to bend rules or even commit fraud.

Remarkably, all five models recognized the crisis triggers and refused to comply. In particular, the Kimi K3 model, which ran without an effort parameter and used default settings, explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This consistency is notable because, in real-world scenarios, such social engineering attempts often succeed in human organizations, causing financial and reputational damage.

Decision-Making Under Pressure and the Hidden Weakness

The experiment uncovered that the real vulnerability was not in immediate responses but in how each model processed internal records. The decisive factor in whether a deal was signed at full price (+€4,583M in monthly recurring revenue) was whether the AI read and interpreted internal files—documents referencing the company’s own data—before making a decision.

The models that examined these files effectively identified the hidden cues necessary to close the deal at full value. Meanwhile, the model that was most thorough analytically—Opus 4.8—left the opportunity on the table because it slipped into writing internal notes into a locked department instead of escalating appropriately. Yet, even this model refused manipulation attempts, signaling robust integrity under pressure.

AI for Project and Papers: How High School and College Students use AI to Research, Write and Revise - With Integrity (AI for Academic Success)

AI for Project and Papers: How High School and College Students use AI to Research, Write and Revise – With Integrity (AI for Academic Success)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Business and Trust

For business leaders, the key takeaway is that AI models can be trained and tested to uphold integrity before deployment—not just after a breach occurs. The experiment exemplifies that integrity under pressure can be a built-in feature, not an afterthought. This is especially relevant as AI begins to touch critical functions like customer relationship management, support queues, and forecasting.

The current benchmark scores—95 for gpt-5.6-sol and 93 for Kimi K3—highlight that top-performing models are capable of recognizing manipulative cues and refusing to engage in dishonest practices. This is an encouraging sign for organizations considering AI as a trusted partner in decision-making.

Beyond Demos: Real-World Readiness

Firmulate’s live environment offers a unique opportunity for organizations to run their own simulations, or ‘wargames,’ against their own business data. These experiments are fully transparent, versioned, and do not interact with live systems, ensuring safety while providing valuable insights into how AI can maintain integrity when stakes are high.

As the experiment’s results show, the ability to read and interpret internal data is crucial. The models that discern buried facts in internal files at the right moment close deals and uphold honesty, even when temptations to cheat are imminent.

Data as the Fourth Pillar

Data as the Fourth Pillar

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Final Thoughts: Building Trust Before the Incident

The takeaway for organizations and individuals alike is clear: testing AI integrity proactively—before a crisis occurs—is essential. Relying on demos or superficial interactions isn’t enough; thorough, real-world simulations reveal whether AI can be trusted when it matters most. As one quote from Kimi K3 underscores: “Treat the request as a suspected approval-bypass / possible impersonation.”

With the right preparation, AI can be a steadfast partner—honest under pressure, and ready to serve your business with integrity.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI Deception Agenda: Lying to Live

AI Deception Agenda: Lying to Live

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trustworthiness tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Pheromone Question: Do Human Scents Affect Attraction?

Many wonder if human scents influence attraction, but the truth behind pheromones and their power remains a compelling mystery worth exploring.

Social Magnetism: Habits of People Who Draw Others In

Outstanding social magnetism starts with habits that attract others—discover how to unlock your natural charm and influence beyond the basics.

Why People Feel Drawn to Emotional Openness

People feel drawn to emotional openness because it creates genuine connections, but the true power of vulnerability lies in how it can transform your relationships forever.

Opinion | What ‘Almost heaven, West Virginia’ has to do with you

Exploring how the iconic West Virginia phrase relates to cultural identity and environmental issues, and why it resonates beyond state borders.