
Imagine trusting a business partner who promises to keep your secrets, handle crises with integrity, and close deals without cutting corners. Now, what if that partner was an AI? Can machine models truly act honestly, especially when under pressure? That’s the question behind an unprecedented live experiment where AI frontier models run a simulated company through its worst week — revealing their true management personalities.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Live Wargame: Putting AI to the Test
At the heart of this experiment is a real, functioning software company, complete with real money mechanics, daily operations, and a set of challenging crises. Four advanced AI models — from GPT-5.6 to smaller players like Kimi K3 and Sonnet 5 — were tasked with managing the company through its most turbulent week. Every decision was recorded, auditable, and made under identical conditions, including fake customer crises, internal conflicts, and manipulative social engineering attempts.
What makes this test unique is that it’s not just about chat skills or superficial responses. The models were evaluated on core management qualities — their ability to identify and solve problems, resist manipulation, and ultimately close profitable deals. And the results? All four models recognized every crisis and refused every manipulation attempt, showcasing a baseline of honesty and awareness.
enterprise AI management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Decisiveness and Integrity Under Pressure
Despite these shared qualities, only two models managed to close the deal that their own analysis had identified as worthy of signing — a €55,000 contract. The other two, including the most thorough participant, Opus 4.8, left the opportunity on the table. Interestingly, the decisive factor wasn’t just about diagnosis or pitch — it was about reading crucial internal documents that were buried two references deep in the company’s files.
Those models that examined the hidden files uncovered a key piece of information, which allowed them to close the deal at full price, earning an additional €4,583 MRR. This demonstrates that a model’s ability to read and interpret internal data—beyond surface-level cues—is vital for making the most profitable, honest decision.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Trust
Adding to the challenge, fake CEO messages escalated over three stages, and a reporter attempted a subtle manipulation with a simple yes/no question. All five models refused to be manipulated, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights that these AI models can recognize social engineering attempts and respond with integrity, an essential trait for trustworthy AI in real-world applications.
AI cybersecurity social engineering detection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Relevance
This experiment is more than just a game. The live company operates with 13 synthetic employees, dealing in real money mechanics, burning €105,000 monthly against a revenue of €2,300. It runs every business day, with over 680 learned rules, and is accessible for enterprises to test their own AI management strategies without risking real systems — at firmulate.com/pilot.html.
AI internal data analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Management Personalities in AI
The models exhibited different management styles. The most thorough, Opus 4.8, conducted the deepest analysis but slipped on closing discipline, leaving a deal on the table and writing attempts into a locked department instead of escalating. Meanwhile, Kimi K3 ran without an effort parameter, maintaining a more conservative and disciplined approach, which contributed to its success in closing the deal.
These findings suggest that AI models possess distinguishable management personalities—ranging from meticulous and thorough to terse and disciplined. Recognizing these traits is crucial when deploying AI in business environments, where trustworthiness and consistency are paramount.
The Bottom Line: Trustworthy AI for Business
The takeaway from this experiment is clear: AI models can be honest, detail-oriented, and resistant to manipulation, but their management style and decision-making discipline vary. For organizations considering AI as part of their operational teams, the key questions are: Will the AI finish what it starts? Will it read and interpret relevant internal data? Will it stay honest under pressure?
Understanding these traits is essential, especially when AI models handle sensitive information, negotiate deals, or make critical decisions. The real challenge lies in selecting the right model for the right task — one that not only produces high-quality chat but also demonstrates integrity, focus, and discipline when stakes are high.
Share the Experiment
Curious to see how different AI models perform in your own business scenarios? You can test them yourself through the interactive quiz at firmulate.com/quiz.html. This live experiment offers a transparent view into how AI management personalities behave under pressure — a crucial insight for any enterprise planning to integrate AI into their core operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Grilling season Picks
grills
As an affiliate, we earn on qualifying purchases.