
Imagine a business so transparent that you can watch its daily struggles, mistakes, and decisions unfold in real time. Now, picture that business running entirely on artificial intelligence, with no human employees but with real money flowing in and out. This is not science fiction; it’s the live experiment of Firmulate, a company that challenges how we think about trust, transparency, and AI’s role in business.
The Unique Experiment in Building Openly
Firmulate operates a small, virtual company with 13 synthetic employees—AI models that make decisions, handle crises, and even negotiate deals. Yet, it’s not just an AI sandbox. The company is real: it burns €105,000 a month, but earns only €2,300 in monthly recurring revenue. Every day, the company’s actions are recorded, versioned, and made publicly accessible at firmulate.com/live.html. This open format offers a rare peek into how AI manages a business under real-world pressures, including crises and ethical dilemmas.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing AI Under Extreme Conditions
Firmulate’s key experiment involves running four advanced AI models through the same difficult week in the company’s life. These models—each representing a different frontier of AI capabilities—face identical scenarios: difficult customer interactions, crises, and even attempts at manipulation. The goal: see if the AI can handle the pressure ethically and effectively.
All four models successfully identified every crisis, refused every manipulation attempt, and maintained operational discipline. Yet, only two of them managed to close a €55,000 deal their own analysis indicated they should have. The other two either left the opportunity on the table or failed to execute the deal, revealing subtle but critical gaps in discipline and focus.
What’s Hidden in the Data?
The most surprising insight isn’t just in the decisions made during crises. It’s in the company’s files—hidden references that even the AI models had to dig deep to find. When some models read these crucial documents, they won the full-price deal, earning an additional +€4,583 in monthly recurring revenue. This highlights that AI’s success often depends on how thoroughly it “reads” the entire information set, not just surface-level interactions.
Ethics and Trust Under AI Management
Another critical aspect tested was AI’s response to social engineering. Fake messages from a CEO, escalating through stages, and even a trick involving a reporter—all designed to test whether the AI would be duped—were systematically refused by every model. Kimi K3, one of the models, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates that the models are not just reactive but are programmed to prioritize security and integrity, even under pressure.
The Human-Like Flaws and Lessons
One of the most thorough models, Opus 4.8, demonstrated that even the best AI can slip up without proper discipline. During the same week, it left an opportunity unexecuted because it failed to escalate a decision properly, instead writing attempts into a restricted department. This shows that even the most advanced AI can mirror human weaknesses, especially under stress or when rules are complex.
The Bigger Picture for Business and Trust
For families and parents, the lessons are surprisingly relevant. Just as a child’s trust depends on consistent, honest behavior, a business depends on reliability and integrity—even when faced with temptations or crises. An AI that can be transparently tested, that refuses manipulation, and that documents its decisions openly—even under extreme pressure—presents a new model for trustworthiness in digital tools.
In a broader sense, firms and enterprises are invited to use this kind of live, transparent testing to evaluate their own AI systems before deployment. The platform offers a way to simulate the worst scenarios, identify weaknesses, and ensure that AI behaves ethically and effectively in real-world situations—just like a family would want a trustworthy caregiver who consistently acts in their best interest.

Firmulate’s open, real-time AI company experiment reveals that trustworthiness, discipline, and thoroughness are critical for AI to replace or support human decision-makers. For families and businesses alike, transparency and rigorous testing are key to building confidence in new AI tools.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html