Executive Overview
For software companies with legacy architectures and established user bases, the arrival of generative AI presented an existential operational dilemma: rebuild from the ground up or risk obsolescence. Few enterprises have navigated this tension at the scale of Atlassian. With a portfolio spanning more than 20 distinct applications—six built in the generative AI era and the remainder reflecting years of legacy development—Atlassian currently serves over 5 million active users relying daily on its AI-native capabilities.
At the helm of this transformation is Sherif Mansour, a 17-year Atlassian veteran who oversees AI across the entire product portfolio and manages the product management craft, a massive organization of 450 product managers (PMs). In his landmark address at the SaaStr AI session, Mansour tackled a problem familiar to every enterprise founder and product leader: the pervasive paralysis of conflicting advice.
The industry playbook for artificial intelligence is riddled with contradictions. Tech pundits demand that companies go entirely headless because conversational interfaces are supposedly all users need. Simultaneously, critics argue that chat is a broken, inefficient user experience that demands dedicated feature development. Industry purists insist on tearing down existing products to rebuild from scratch, warning against simply "bolting on" AI, while others argue that every team member should become a hands-on AI builder.
Rather than adhering to rigid theoretical dogma, Atlassian tested these conflicting paradigms at massive scale with real teams and live users. The results challenge conventional wisdom across three critical dimensions: the necessity of conversational interfaces, the strategic viability of bolting AI onto legacy workflows, and the limits of collapsing product management, design, and engineering into a single "AI builder" role. This report provides an investigative breakdown of Atlassian’s strategic decisions, the metrics driving their R&D pipeline, and the operational realities of deploying AI across a sprawling enterprise portfolio.
Detailed Chronology & Strategic Decisions
Decision 1: The Calculated Bet on Chat and the DOS Analogy
Two years ago, integrating Rovo chat natively into core products like Jira and Confluence was deeply controversial within Atlassian’s internal teams.
The case against chat was compelling: conversational interfaces are rarely an ideal end state for complex enterprise tasks. Forcing users to fill out complex, structured forms or execute multi-step database queries through a conversational prompt is often the least efficient user experience imaginable. Furthermore, users already had access to standalone external intelligence tools like ChatGPT and Anthropic’s Claude.
Conversely, the case for chat was rooted in uncertainty. Because consumer and enterprise behavior in the age of generative AI remained entirely unmapped, product teams needed an open-ended interface to observe how users naturally attempted to solve problems.
The debate was ultimately settled by practical engineering necessities. Atlassian had to build a centralized chat backend anyway to power localized AI features across its broader product portfolio. Once that infrastructure existed, the company shipped the conversational interface to more than 20 applications simply to observe usage patterns and gather empirical data.
To explain the enduring utility of chat, Mansour relies on a historical analogy: the Disk Operating System (DOS). In the early days of personal computing, the command line served as the universal interface for the operating system. Users executed spreadsheets, word processing, and games within a single text interface. Over time, as standard use cases matured, they spun off into dedicated applications featuring rich, graphical user interfaces. Yet, the command line never truly disappeared; it continued to serve the long tail of irregular, edge-case interactions.
Chat plays an identical role in the modern AI stack. It absorbs unlimited, unstructured use cases and acts as an observational testing ground, revealing precisely which user interactions command the development of a dedicated, purpose-built graphical user interface.
From Whiteboard Prompts to Multi-Agent Workflows
Atlassian tracked this evolutionary pathway directly within Confluence whiteboards. By analyzing what users naturally typed into open-ended prompts, the product team identified recurring usage patterns that quickly transformed from simple text commands into native product features, structured workflows, and eventually, callable agent tools.
Crucially, Atlassian realized that every feature engineered for human users had to simultaneously be exposed as an API tool for autonomous agents. Capabilities such as grouping sticky notes into strategic themes or converting whiteboard brainstorms directly into software backlogs are now standard skills that Rovo agents can invoke autonomously. By supporting Model Context Protocol (MCP) and Atlassian’s command-line interface, these capabilities can be integrated into external developer environments like Claude Code and Cursor. Consequently, Atlassian’s proprietary software features provide utility whether the user is actively inside an Atlassian web application or working externally.
This empirical reality dismantled Atlassian’s initial assumption that chat was merely a temporary onboarding ramp. Millions of users engage with Rovo chat daily, even while simultaneously leveraging competing systems like Gemini and Claude. Enterprise users gravitate toward the artificial intelligence embedded closest to their active workspace.
Mansour’s actionable recommendation for enterprise leaders is direct: if your product retains daily active users, ship conversational chat. Use open-ended user behavior to dictate your feature roadmap rather than relying on internal speculation.
Decision 2: The Pragmatic Reality of Bolting AI Onto Legacy Workflows
Standard software engineering advice strongly warns against "bolting on" artificial intelligence to legacy products, dictating that companies must completely reimagine architectures from scratch. Atlassian intentionally chose to violate this rule.
To understand this decision, one must examine the true nature of Jira. New product managers joining Atlassian frequently operate under the assumption that Jira is exclusively a tool for software developers. In practice, the vast majority of Jira users are not software engineers. It functions as a generalized workflow engine across diverse business operations. If a customer complains about a rideshare trip, or disputes a telecommunications billing error, those administrative inquiries travel through customized Jira workflows.
Two and a half years ago, before market trajectories for enterprise AI were clear, Atlassian took its existing Jira automation designer and added a single new configuration box: the ability to invoke a Rovo agent at a specific step in the sequence. It was a literal, unpolished bolt-on to a legacy architectural foundation.
Enterprise customers embraced the capability and rapidly pushed it far beyond initial design expectations. Organizations introduced conditional logic, branching pathways, and chained autonomous agents together. In one prominent customer pattern, an incoming support ticket triggers an automated agent; that agent calls Canva to programmatically generate marketing creative, and a secondary social media agent reviews the assets before scheduling them for publication. Customers simply automated the human-run operational workflows they had already established inside Jira over the previous decade.
Mansour illustrates this philosophy with a domestic analogy: when his family purchased their home, they deliberately lived with the outdated, existing kitchen before initiating a renovation. Living within the legacy space revealed precisely which functional bottlenecks warranted capital expenditure and structural redesign.
If an enterprise possesses active users and established operational workflows, those workflows constitute a core business asset that should be systematically evolved. Rewriting products from scratch is a strategy best reserved for entirely new product lines—an approach Atlassian executed concurrently with its six modern, AI-native applications.
The Core Principle: Human-Agent Parity
Approximately six months ago, Atlassian codified a guiding principle derived from its early bolt-on experiments: Anything you can accomplish with a human team member, you can accomplish with an autonomous agent.
Every core requirement the organization had previously solved to make human collaboration functional—tools, operational context, explicit goals, accountability, and real-time visibility into peer activities—applied identically to artificial agents. Once enterprise users began deploying networks of agents concurrently, high-level planning across both human and artificial team members emerged as a central operational challenge.
This realization prompted Atlassian to evaluate its product portfolio item by item, establishing a strict internal design rule: whenever an internal team demonstrates an agent capability that a human user could not or would not execute within the product, designers must rigorously audit the rationale. While legitimate technical exceptions exist, nine times out of ten, the autonomous agent should be engineered to mirror the interaction patterns already established for human operators.
Decision 3: The Limits of the "AI Builder" and the Collapse of Specialization
As generative coding assistants matured, industry observers predicted the immediate convergence of product management, user experience design, and software engineering into a singular role: the "AI builder." With an organization comprising 450 PMs and roughly 800 designers, Atlassian needed to test whether this structural collapse was viable at enterprise scale.
To evaluate the hypothesis, Atlassian launched ten dedicated software engineering projects staffed entirely with hand-picked AI builders—engineers alongside PMs and designers capable of writing production code using generative tools across both legacy and modern codebases.
The initial weeks showed explosive velocity. Teams shipped features at unprecedented speeds. However, after a period ranging from several weeks to a few months, momentum ground to a severe halt.
Mansour describes the operational breakdown succinctly: everyone was aggressively rowing, but nobody was steering. Without dedicated leadership making hard strategic calls on product direction, target personas, and scope, the teams began spinning in circles. Crucially, the product managers and designers voluntarily drifted back to their traditional functional roles. Supplying deep customer context, prioritizing backlogs, and making high-stakes trade-offs proved to be the most valuable contributions they could make to unblock their engineering counterparts.
The underlying operational ratios explain this friction. Atlassian traditionally maintains a ratio of roughly one PM and designer per ten engineers on product teams, stretching to one per twenty on platform initiatives. As generative coding tools empower engineers to output raw code at multiples of their previous capacity, the effective operational ratio expands to one PM for every thirty or forty engineers. Engineers rapidly outpace product clarity, returning to management constantly to ask what should be built next and whether current assumptions are strategically sound. A product manager spending their workday "vibe coding" is fundamentally unavailable to provide that strategic governance.
While Atlassian continues to carve out specialized spaces for technical builders, Mansour confirms that this is not a viable paradigm for the majority of product managers or designers in a mid-sized or large enterprise. In an early-stage startup, everyone can function as a generalized builder; in a scaled corporate organization, structural specialization remains essential.
Supporting Context, Metrics, & R&D Evolution
The Inversion of the Hiring Pipeline: Juniors, Seniors, and the Squeezed Middle
The enterprise discourse around technical recruitment has long been dominated by a singular executive directive: cease hiring junior engineers, eliminate entry-level pipelines, and invest capital in senior engineering talent amplified by AI productivity tools. The prevailing critique claims that junior engineers produce low-quality code ("slop") while requiring a multi-year payback period, whereas senior engineers leverage AI to multiply their enterprise output.
Atlassian’s empirical research and hiring data forced a radical departure from this conventional wisdom. The company’s recruitment pipeline has structurally inverted: Atlassian now indexes heavily on both ends of the seniority spectrum—new graduates and deeply seasoned senior staff—while compressing the middle tier.
Why New Graduates Matter in the GenAI Era
The rationale for retaining and expanding the junior engineering pipeline is rooted in the cognitive friction of enterprise software development. Atlassian employs thousands of R&D professionals who have spent ten, twenty, or thirty years mastering specific enterprise development methodologies. Unlearning established habits is cognitively more difficult than learning novel workflows.
New university graduates, by contrast, possess no legacy habits to unmar. To illustrate this generational shift, Mansour notes that his twelve-year-old child bypassed traditional syntax education entirely, moving directly from visual programming languages to generative coding tools to construct an application for commercial app stores on Replit. The next generation views generative AI as the native syntax of software development.
Atlassian’s internal R&D research validates this operational thesis:
- Junior engineers are up to 38% more likely to experiment with generative AI tools.
- They utilize nearly twice as many distinct AI tools in their daily workflows compared to mid-level peers.
- Conversely, junior engineers frequently produce unrefined output and do not self-report feeling exceptionally productive.
- Senior engineers adopted AI tools at a slower initial rate, but demonstrated superior capacity for identifying architectural flaws ("slop") and extracting high-value outputs from language models, drawing on decades of experience managing human juniors.
To bridge this generational divide, Atlassian orchestrates organization-wide AI Builder Weeks every few months. For four synchronized days, the entire R&D division pauses standard feature delivery. Junior engineers host technical sessions demonstrating cutting-edge AI tools and prompt methodologies, while senior staff lead sessions focused on quality control, architectural integrity, and differentiating proprietary AI applications from generic wrappers. Atlassian has open-sourced the program structure, operational schedules, and recorded instructional sessions for the broader tech community.
Mansour’s ultimate synthesis on enterprise human capital is clear: hire individual AI builders when available, but architect your organization to cultivate 10x teams rather than relying on isolated 10x individuals.
Future Outlook
As Atlassian looks toward the next phase of enterprise software development, its operational experiments offer a blueprint for navigating the transition from legacy architectures to agentic workflows. The roadmap ahead highlights several key focal points for enterprise product leaders:
- The Maturation of Agentic Interoperability: As tools like the Model Context Protocol (MCP) gain broader enterprise adoption, the boundary between proprietary application silos will continue to dissolve. Software must be architected so that internal business logic is natively exposed as callable tools for autonomous agents operating across external third-party environments.
- Rebalancing Organizational Ratios: Enterprises must solve the structural imbalance created by AI-accelerated engineering velocity. Organizations must determine how product management and design functions scale their strategic guidance to prevent engineering teams from rapidly building sophisticated products in the wrong strategic directions.
- Continuous Architectural Evolution: The false dichotomy between total architectural rewrites and legacy stagnation must be abandoned. Enterprises with established user bases should embrace pragmatic, iterative integration—bolting on automation steps to validate market demand before committing capital to deep structural overhauls.
By systematically testing opposing industry dogmas against empirical customer behavior, Atlassian has demonstrated that successful generative AI integration is not about wholesale abandonment of the past, but about methodically bridging legacy infrastructure with agentic futures.
