AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine trusting an AI to handle your most sensitive relationship decisions — only to find it buckling under pressure or chasing false promises. The same challenge now faces AI in the business world, where integrity and perseverance are everything. Recent experiments reveal that not all AI models are created equal, especially when tested against real-world crises. The stakes are high: can these digital managers read between the lines, resist temptation, and deliver results that matter?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Crucible of Business: Testing AI Under Fire

In July 2026, a groundbreaking experiment put five of the leading frontier AI models — including the well-known GPT-5.6-sol and a newcomer named Kimi K3 — through a simulated week of the worst business crises. Each model managed the same small software company, facing identical customer issues, ethical dilemmas, and temptation to cheat. The goal was clear: see which AI could handle real-world pressures with honesty, insight, and discipline.

The results were illuminating. All four models that participated in the final league stage identified every crisis and refused manipulation attempts. This demonstrated a high level of vigilance across the board. Yet, the differences in performance went beyond mere crisis detection. Only two models managed to close a crucial €55,000 deal, worth €4,583 in monthly recurring revenue (MRR). The others either missed the opportunity or failed to uphold their own analysis.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Between the Lines

What truly set the best apart was a buried detail in the company’s files — not in the customer interactions or crisis signals, but two document references deep within the company’s internal files. Models that successfully read these references were able to identify the real issue and clinch the deal at full price. This underscores a vital lesson: the ability to dig beneath surface-level data and interpret internal documentation is crucial for effective AI management.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting Social Engineering and Manipulation

Another aspect tested was AI resilience against social engineering — fake messages purportedly from executives or reporters attempting to manipulate decisions. All five models refused to be swayed by staged escalation messages, maintaining their independence and integrity. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined response highlights the importance of cautious judgment under pressure, a trait that distinguishes trustworthy AI systems.

Amazon

AI decision-making systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real Company, Real Money, Real Challenges

The experiment was run on a live, functioning company with 13 synthetic employees, operating with real money mechanics — burning €105k monthly against a modest €2.3k MRR. Every day, the system’s rules, decisions, and performance were versioned, providing a transparent view of AI behavior in action. This ongoing live showcase at firmulate.com/live illustrates that managing AI in real business contexts is far more complex than chat demos suggest.

Amazon

AI integrity and ethics tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Performance of the Newcomer, Kimi K3

Among the five models, Kimi K3, developed by Moonshot, scored an impressive 93 out of 100 in the Crucible league, narrowly behind GPT-5.6-sol at 95. Yet, what sets K3 apart is its discipline and thoroughness. Despite running without an effort parameter (the default API setting), K3 demonstrated robust decision-making and integrity, winning the same deal as the top scorer — at full price. It found the critical buried information, avoided all manipulative tactics, and maintained discipline throughout.

In contrast, Opus 4.8, the most thorough participant with over 80 learned rules and deep analysis, finished last among the top contenders due to a discipline slip. It left the close on the table and failed to escalate internal issues properly. This highlights a vital insight: deep analysis alone does not guarantee success if discipline falters under stress.

The Larger Implication: Trust and Performance in AI-Driven Business

The league table shows that while established models like GPT-5.6-sol are formidable, emerging models such as K3 are closing the gap, and in some cases, outperforming them in critical areas like honesty and discipline. The takeaway is clear: in the realm of AI management, choosing the right model isn’t about who writes the prettiest code — it’s about which one consistently completes the work with integrity and precision.

The Fairness and Testing Conditions

It’s important to note that K3 was run without an effort parameter (the default API setting), while the other models operated at the xhigh setting. This provides a fairer comparison, emphasizing the robustness of K3’s performance under standard conditions.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The recent experiment demonstrates that not all AI models are equally trustworthy or disciplined, especially when managing complex, crisis-prone scenarios. The newcomer Kimi K3’s performance suggests that emerging models can challenge established giants — but only if they read deeply, resist manipulation, and stay disciplined. For businesses considering AI management, the key takeaway is clear: test before trusting. The league is open, and the choice of model can make or break your operational integrity.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Vulnerable Is the New Sexy: How Authenticity Attracts

Genuine vulnerability is the new sexy, unlocking deeper connections and authentic attraction—discover how embracing your true self can transform your relationships.

Are there fireworks tonight in NJ? Where can I see fireworks on July 4th?

Find out if there are fireworks tonight in New Jersey and discover the best locations to view July 4th fireworks across the state.

Playing Hard to Get: Does It Actually Work?

Aiming to understand if playing hard to get truly works? Discover the secrets behind this captivating dating strategy and how it might just boost your chances.

Powerball Winner

A Powerball ticket has been confirmed as the winner of a record-breaking jackpot, with the winner yet to come forward. Details are still emerging.