The Transparency Paradox: Why B2B AI Vendors Must Publish Their Deepest Competitive Evals to Win

Share
The Transparency Paradox: Why B2B AI Vendors Must Publish Their Deepest Competitive Evals to Win

Executive Overview

In the hyper-competitive landscape of business-to-business (B2B) software, a quiet revolution is underway. For decades, vendor differentiation relied on static marketing assets: glossy analyst quadrants, self-reported feature checklists, G2 grid dominance, and PowerPoint slides boasting unverified performance metrics like "3x faster resolution." Buyers grew accustomed to filtering out the noise, treating vendor claims with a healthy dose of skepticism.

However, the proliferation of generative artificial intelligence and autonomous AI agents has rendered traditional software benchmarking obsolete. Unlike traditional software—where a CRM database executes deterministic code predictably every time—AI agents are dynamic, probabilistic systems. They can deliver stellar performance during a Monday demo, hallucinate or fail on a Thursday production run, and degrade silently following an unannounced underlying model update.

Enter the age of the competitive evaluation (eval).

Recently, Gorgias—a $100M ARR leader in ecommerce customer experience (CX) software, backed by the SaaStrFund seed round—shattered industry convention by publicly open-sourcing its entire AI agent evaluation harness. Testing 8,356 live customer conversations across 18 competing market vendors, Gorgias published a brutally honest comparative analysis. Crucially, the report did not merely crown Gorgias the victor; it transparently highlighted metrics where rival tools, such as Yuma and Envive, decisively outperformed them.

This bold maneuver marks a turning point for the enterprise software ecosystem. As AI agents increasingly assume the responsibility of building shortlists and making purchasing recommendations, B2B vendors face a stark choice: embrace radical, transparent evals or lose the trust of modern buyers and autonomous agents alike.


Detailed Chronology: How Gorgias Reset the Bar for AI Benchmarking

The genesis of the Gorgias public AI benchmark reflects a deliberate shift in go-to-market strategy, driven by an industry-wide need for verifiable truth.

The Seed of Radical Openness

Long before the public launch of the benchmark, investors and industry veterans—including SaaStr’s Jason Lemkin—pushed portfolio companies to move beyond surface-level feature comparisons. In the burgeoning agentic AI economy, marketing claims carry zero weight if they cannot be mathematically reproduced and verified. Gorgias, realizing that roughly 80% of its ~$100M ARR was increasingly tied to its AI support layer, accepted the challenge.

Building the Harness

Rather than relying on synthetic datasets or controlled laboratory environments, Gorgias engineering teams constructed a comprehensive testing harness designed to evaluate real-world friction.

  • The Scale: The team fed 8,356 live ecommerce customer service conversations into the testing framework.
  • The Competition: The harness evaluated 18 distinct vendor agents simultaneously against the exact same baseline data.
  • The Rubric: Every response was measured against a stringent, version-controlled rubric, scoring dimensions such as resolution accuracy, multi-turn reasoning, and latency (with competitors like Envive clocking blistering response speeds around 7.9 seconds).

The Public Release

When Gorgias officially published the benchmark report—supported by an open-source GitHub repository containing the entire evaluation harness—it sent ripples through the SaaS community. By making the code, scoring weights, and raw data completely checkable, Gorgias bypassed traditional gating mechanisms. Even though the company secured the #1 position in overall support efficacy, it prominently displayed where competitors led in automation success rates and speed. The message to the market was clear: Here is the raw data. Verify it yourself.


Supporting Context & Metrics: Why Traditional Software Benchmarks Fail AI

To understand the strategic brilliance of Gorgias’s move, one must examine why traditional benchmarking mechanisms are fundamentally broken when applied to artificial intelligence.

The Anatomy of an "Eval" vs. A "Benchmark"

In legacy enterprise software, a benchmark evaluates static capabilities: Does the platform support single sign-on (SSO)? Does it integrate with Shopify? How many API calls can it handle per minute? These attributes change slowly and predictably.

An AI evaluation, conversely, measures behavioral reliability:

  1. Real-Time Responsiveness: How does the agent handle ambiguous human emotions, sarcasm, or complex refund logic?
  2. Deterministic vs. Probabilistic Drift: Does the agent maintain high performance across thousands of edge cases, or does it fail under multi-turn conversation pressure?
  3. Continuous Variance: Because underlying foundational models (like GPT-4o, Claude 3.5 Sonnet, or custom fine-tuned weights) shift frequently, an AI agent’s performance is a moving target.

By executing 8,356 live conversations, Gorgias effectively simulated months of frontline customer support interactions in a matter of hours, exposing the performance deltas that no vendor-supplied demo could ever reveal.

The Psychology of Modern B2B Buying

Modern procurement leaders are exhausted by spin. When a vendor publishes a report claiming absolute supremacy across every conceivable metric, enterprise buyers instinctively discount the data.

Conversely, when a market leader publishes an eval that explicitly showcases competitor victories, a psychological shift occurs. Radical transparency breeds credibility. If Gorgias admits that a rival vendor excels in raw response velocity or niche automation workflows, the buyer is far more likely to trust Gorgias’s claimed victories in overall support satisfaction.

Furthermore, this approach fundamentally streamlines enterprise due diligence. Historically, an ecommerce brand evaluating support automation would sit through weeks of sales pitches, conduct shallow three-vendor demos, and run an error-prone internal pilot. By open-sourcing a rigorous, 18-vendor evaluation framework, Gorgias essentially performed the industry’s heavy lifting, positioning its report as the definitive industry standard.


Official Statements & Industry Perspectives

The release of the Gorgias benchmark has sparked intense debate and commentary across the venture capital and SaaS leadership communities.

Industry leaders have been quick to praise the strategy as a masterclass in modern positioning. Jason Lemkin, founder of SaaStr and early investor through the SaaStrFund, encapsulated the sentiment shared by many market observers on social media:

"Everyone should publish the deepest, most direct competitive evals they can. Will there always be some bias? Yes. But Gorgias did it the right way in ecomm CX: Open-sourced the whole harness, tested 8,356 live conversations across 18 vendors, and while taking #1 in support, proudly showed where competitors like Yuma excelled."

This perspective highlights a crucial nuance: bias in evaluations is inevitable because the creator of the rubric ultimately defines the scoring weights. However, rather than masking this inherent bias behind closed doors, Gorgias exposed it. By publishing the scoring weights and open-sourcing the evaluation code on GitHub, they invited the developer and buyer communities to inspect, critique, and even run the harness themselves.

Tech industry analysts have similarly noted that this transparency serves as a powerful internal forcing function. When competitive performance metrics—such as Envive’s 7.9-second latency or rival resolution rates—are published publicly alongside internal figures and updated weekly, complacency becomes impossible. Engineering and product teams no longer rely on obscured internal telemetry; they track against a living, breathing public scoreboard.


Future Outlook: The Rise of Agentic Procurement

Looking ahead, the publication of deep, open-source competitive evals points toward a profound transformation in how enterprise software will be bought and sold over the next decade: the rise of agentic procurement.

When Autonomous Agents Do the Shortlisting

We are rapidly entering an era where human buyers are no longer the primary consumers of B2B marketing collateral. Increasingly, enterprise software discovery and vendor shortlisting are delegated to autonomous research agents—AI systems tasked by procurement teams to evaluate market options.

An LLM-driven research agent cannot read a gated PDF marketing brochure or be swayed by a flashy brand video. It requires structured, checkable data: versioned rubrics, open-source repositories, explicit scoring weights, and reproducible evaluation code.

Currently, friction remains. Notably, Gorgias initially blocked automated readers via its website’s robots.txt file on the live report—a minor tactical oversight that underscores how legacy web governance habits clash with the needs of agentic discovery. As the ecosystem matures, vendors that seamlessly expose their evaluations in machine-readable formats will dominate AI-generated vendor shortlists. Those relying on gated, proprietary PDF whitepapers will find themselves invisible to the autonomous agents driving enterprise decisions.

The Blueprint for Future AI Vendors

For startups and scale-ups entering the AI arena, the Gorgias playbook provides a clear template for establishing market authority:

  1. Test on Production Reality: Ditch synthetic benchmarks and evaluate agents against thousands of authentic, messy, real-world customer interactions.
  2. Name Names and Share the Losses: Do not shy away from competitor strengths. Highlighting where rivals win validates your own legitimate strengths.
  3. Open-Source the Harness: Publish the testing code, the exact prompts, and the grading rubrics. Make the evaluation entirely reproducible.
  4. Version Control Your Rubrics: Ensure scoring metrics are immutable and transparent so the market knows you aren’t secretly moving the goalposts to hide performance gaps.
  5. Optimize for Machine Discovery: Ensure your evaluation data is fully accessible to the automated agents and LLMs that modern buyers rely on for initial vendor filtering.

Ultimately, the era of smoke-and-mirrors software marketing is drawing to a close. In an AI-first world, transparency is no longer just a moral or ethical choice—it is the ultimate competitive moat. Vendors that boldly open their books, publish their evals, and invite public scrutiny will earn the trust of both human buyers and the intelligent agents of tomorrow.

Did you find this story helpful?

Share it with your friends and colleagues on social media.

Share

Leave a Comment

Your email address will not be published. Required fields are marked *