Navigating the Generative AI Era: Atlassian’s Blueprint for Scaling AI Across Legacy and Native Products

Share
Navigating the Generative AI Era: Atlassian’s Blueprint for Scaling AI Across Legacy and Native Products

Executive Overview

The generative artificial intelligence revolution has presented B2B software enterprises with a perplexing paradox. For every expert advocating for headless architecture and chat-first interfaces, an equally adamant voice demands deep native workflows and traditional graphical user interfaces. Similarly, industry dogma insists that companies must never simply "bolt on" AI features to legacy infrastructure, yet starting entirely from scratch risks alienating millions of incumbent users.

How do you reconcile these conflicting directives when managing an expansive software portfolio that spans nearly two decades?

Sherif Mansour, Head of AI at Atlassian and leader of a massive product management organization overseeing more than 450 PMs, addressed this exact challenge during his widely discussed SaaStr AI session. Atlassian’s real-world laboratory offers a masterclass in enterprise AI deployment. Managing a portfolio of more than 20 distinct applications—six built natively in the generative AI era, and the remainder representing years of legacy architecture—Atlassian now serves over 5 million active users who regularly leverage its integrated AI capabilities.

Rather than adhering to rigid, dogmatic rules, Atlassian tested every prevailing AI hypothesis at scale with real teams and real users. The results yielded three foundational product decisions, a radical shift in hiring strategy, and invaluable insights into the realities of modern software engineering. This report details Atlassian’s journey, exploring how enterprise giants can successfully navigate the integration of artificial intelligence without destabilizing their core business.


Detailed Chronology: Atlassian’s Three Defining AI Decisions

When the generative AI wave crested, Atlassian was forced to make high-stakes decisions under conditions of deep market uncertainty. Their approach relied on empirical testing, pragmatic compromise, and a willingness to embrace tactics that conventional wisdom outright forbade.

Decision 1: Embracing the Chat Interface as the Enterprise "DOS"

Two years ago, integrating Rovo chat directly into Jira, Confluence, and the broader Atlassian application ecosystem was a subject of intense internal controversy.

  • The Case Against Chat: Critics argued that chat was merely a transitional state rather than a permanent destination. For complex enterprise tasks, a conversational interface is often the worst possible user experience. Asking a user to fill out a structured, multi-field form purely through chat, for instance, introduces unnecessary friction and error.
  • The Case For Chat: Proponents countered that user intent in the generative AI era was entirely unchartered. Nobody truly knew what enterprise users would attempt to do with AI tools. Providing an open, flexible text box allowed the company to observe real-world behavior directly.

Ultimately, a mundane engineering reality tipped the scales: Atlassian had to build a robust chat backend anyway to power localized AI features across its product suite. Once that infrastructure existed, they deployed the chat interface across all 20-plus applications to analyze user patterns.

Mansour conceptualizes chat through an insightful historical analogy: the Disk Operating System (DOS) command line. In the early days of personal computing, the command line served as the universal interface for the operating system. Users executed spreadsheets, word processors, and games all from a single text prompt. Over time, as common use cases crystallized, those functions migrated into dedicated applications with specialized graphical user interfaces.

However, the command line never truly disappeared; it continued to efficiently handle the long tail of complex, ad-hoc tasks. Chat plays an identical role in the modern AI ecosystem. It accommodates unlimited, unbounded use cases while illuminating which specific workflows deserve dedicated, purpose-built interfaces.

From Whiteboard Prompt to Omnichannel Agent Tool

To illustrate this evolution, Mansour pointed to Confluence whiteboards. By analyzing what users naturally typed into prompts, Atlassian engineers identified three dominant usage patterns. Surprisingly, these human-centric features revealed an immediate secondary utility: every function built for a human user had to simultaneously be exposed as a tool for autonomous AI agents.

For example, a prompt designed to "group-into-themes" on a whiteboard evolved into a core skill that Rovo agents could call programmatically. Similarly, "whiteboard-to-backlog" translation became an automated pipeline. By exposing these capabilities through the Model Context Protocol (MCP) or Atlassian’s Command Line Interface (CLI), customers could integrate them seamlessly into external developer environments like Claude Code and Cursor. Consequently, Atlassian’s native features generate value whether the end user is actively inside an Atlassian application or working externally.

Decision 2: The Pragmatic "Bolt-On" Approach to Jira Workflows

Conventional product management orthodoxy dictates that software companies must never bolt AI features onto an existing product, demanding instead that legacy systems be completely reimagined and rebuilt from scratch. Atlassian deliberately violated this rule 2.5 years ago.

To understand why, one must look at Jira’s true nature. While new product managers frequently arrive at Atlassian under the assumption that Jira is exclusively a developer-centric tool, the empirical reality is vastly different. The majority of Jira users are not software engineers. Jira functions fundamentally as a universal workflow engine for every conceivable enterprise team. Mansour notes that a customer complaining about a rideshare journey or disputing a telecommunications bill is often having their grievance processed through an underlying Jira workflow.

Recognizing this, Atlassian took its existing Jira automation designer and executed a literal bolt-on: they added a single new configuration box allowing users to invoke a Rovo agent at any designated step in a workflow. This was shipped rapidly because leadership acknowledged they could not predict the trajectory of the market and needed to learn at high velocity.

Customers immediately pushed this capability far beyond initial expectations. Enterprise users began introducing complex branching logic, conditional dependencies, and chained agent architectures. In one prominent pattern observed by the team, an incoming support ticket triggers an automated agent; that agent calls Canva to dynamically generate marketing assets, and a secondary social media agent simultaneously reads those assets and schedules posts across platforms. Customers simply took their existing human-run workflows and automated them using familiar paradigms.

Mansour uses a residential analogy: when his family purchased their home, they deliberately lived with the existing, outdated kitchen for a period before launching a renovation. Living within the old space revealed precisely which operational bottlenecks were genuinely worth solving. They modified pieces iteratively before ultimately gutting the room.

For established products with active users and deeply ingrained workflows, the legacy system is an invaluable asset to be evolved, not discarded. "From-scratch" development is appropriately reserved for greenfield products—a strategy Atlassian pursued in parallel with its six AI-native applications.

Decision 3: The Short-Lived Rise and Fall of the "AI Builder" Squads

With the advent of advanced code-generation tools, a popular industry narrative emerged: the traditional boundaries separating product management, design, and engineering would permanently collapse into a single, unified "AI builder" role. Managing 450 product managers and approximately 800 designers, Atlassian needed to test this hypothesis empirically.

The company stood up roughly 10 dedicated software projects staffed exclusively by hand-picked AI builders—engineers paired with PMs and designers capable of writing code directly via AI across both new and legacy codebases.

  • Initial Phase: The first few weeks were characterized by staggering velocity. Teams shipped features at an unprecedented rate.
  • The Stagnation Point: After several weeks to a few months, however, progress ground to a severe halt.

Mansour describes the phenomenon succinctly: everyone was actively rowing, but no one was steering. Because all participants were immersed in execution, critical decisions regarding product direction, prioritization, and strategic scope languished. Teams began spinning in circles. Ultimately, the product managers and designers voluntarily drifted back to their traditional roles because supplying crucial customer context, defining requirements, and making hard strategic trade-offs proved infinitely more valuable in unblocking engineering output than writing code.

The structural ratios within Atlassian help explain this reality. Product teams typically operate at a ratio of roughly one PM and designer for every 10 engineers, while platform teams run at one to 20. When AI tools empower engineers to write code at triple or quadruple their previous speed, those ratios effectively stretch to 1-to-30 or 1-to-40. Engineers rapidly deplete their backlog and return demanding clarity on the next objective. A product manager spending their day "vibe coding" is unavailable to provide that strategic guidance.

Consequently, while Atlassian continues to embrace technical product managers and specialized builders, the enterprise concluded that the universal collapse of core product roles into a monolithic AI builder is untenable at scale.


Supporting Context & Metrics: Reshaping the R&D Hiring Pipeline

Perhaps the most counterintuitive operational shift at Atlassian occurred within its human resources and talent acquisition pipeline.

Flipping the Pipeline: Juniors, Seniors, and the Hollowed Middle

Executive leadership across the technology sector frequently echoes a uniform directive: Stop hiring junior developers and instead hire a single senior engineer augmented by generative AI. The prevailing rationale is that junior engineers produce excessive "slop" (low-quality, unmaintainable code), seniors accomplish exponentially more work independently, and junior onboarding requires a multi-year investment before yielding a positive return.

Atlassian deliberately tested this thesis and arrived at the exact opposite conclusion. Historically, the company’s hiring pipeline favored a modest intake of junior talent heavily balanced by a large volume of mid-level and senior engineers. Today, that pipeline has inverted: Atlassian weights its hiring toward both ends of the seniority spectrum—juniors and seniors—while significantly reducing mid-level recruitment.

The Unlearning Dilemma

The reasoning is rooted in cognitive science and organizational psychology. Atlassian employs thousands of research and development professionals who must actively unlearn operational habits ingrained over 10, 20, or 30 years of legacy software development. In enterprise environments, unlearning established muscle memory is demonstrably harder than learning entirely new paradigms.

New graduates, by contrast, carry zero legacy baggage. Mansour notes that his 12-year-old child bypassed traditional syntax education entirely, transitioning straight from visual coding languages to vibe coding on platforms like Replit to build production-ready applications. To the next generation, generative AI interaction is the default definition of software development.

Internal Research Findings

Atlassian’s rigorous internal research validates this generational divergence:

  • Junior developers are up to 38% more likelier to actively utilize generative AI tools.
  • Juniors are nearly twice as likely to experiment with novel workflows and utilize a broader stack of utilities.
  • Conversely, while juniors produce higher volumes of initial syntactic "slop" and often report lower subjective feelings of personal productivity due to high experimentation fatigue, senior engineers exhibit slower initial adoption rates.
  • However, senior engineers excel dramatically at quality control, identifying structural code slop, and extracting superior, deterministic outputs from AI models—skills honed through years of mentoring human junior developers.

To bridge this divide, Atlassian instituted AI Builder Week across its entire R&D organization every few months. For four synchronized days, all standard product development pauses. Junior developers conduct workshops on emerging tools and experimental techniques, while senior engineers lead masterclasses on architectural governance, code quality control, and building differentiated AI features rather than generic wrappers. Atlassian has publicly open-sourced the program structure, schedules, and session recordings to aid the broader software community.


Official Statements & Strategic Principles

Out of Atlassian’s extensive experimentation emerged a universal architectural rule and a clarifying governance principle that now governs their entire product portfolio.

The Core Portfolio Principle: Human Primitives Equal Agent Primitives

Approximately six months ago, Atlassian’s product and engineering leadership codified a foundational operational rule: Any capability required to manage a human team is structurally identical to the capability required to manage an autonomous AI agent.

When systematically evaluated, the core primitives required for effective collaboration are the same across both carbon- and silicon-based teammates:

  1. Tools: Both humans and agents require functional applications to execute tasks.
  2. Context: Both entities demand historical data, institutional knowledge, and situational awareness.
  3. Goals and Accountability: Both require explicit objectives, quantifiable success metrics, and performance tracking.
  4. Visibility: Both require real-time transparency into what teammates and adjacent systems are actively processing.

When thousands of organizational agents begin executing parallel workflows within an enterprise, cross-team planning becomes exponentially complex. Consequently, Atlassian audited its product portfolio item by item to ensure that agent management systems mirrored existing human management workflows.

The Design Review Rule

This realization generated a strict internal governance guideline for product design reviews: When a product team demonstrates a capability executed by an AI agent that a human user fundamentally could not—or would not—perform within the application, reviewers must immediately interrogate the rationale.

While legitimate technical justifications occasionally exist, Mansour notes that nine times out of ten, the agent should be re-architected to follow the exact behavioral and operational patterns already established for human users.


Future Outlook: Key Takeaways for Enterprise Founders and Product Leaders

For technology executives, founders, and product leaders grappling with the integration of generative AI into established software products, Sherif Mansour’s roadmap provides definitive strategic guidance:

  • Embrace Chat as a Discovery Engine: Do not dismiss conversational interfaces prematurely. Even if chat is not the ultimate enterprise UI, it serves as an indispensable command-line interface that reveals the long tail of user intent and validates which features warrant dedicated graphical interfaces.
  • Evolve Rather Than Gut Legacy Assets: If your software possesses active daily users and deeply embedded enterprise workflows, treat those workflows as core assets. Bolt on automation capabilities iteratively to observe real-world usage patterns before committing to extensive architectural overhauls.
  • Reject the Monolithic "AI Builder" Fallacy: While engineering velocity will accelerate, product management and strategic design cannot be eliminated. Scaling requires dedicated human oversight to ensure teams are actively steered toward meaningful customer problems rather than spinning in circles.
  • Rebalance Talent Pipelines: Do not abandon junior talent in pursuit of senior-only AI augmentation. Combine the unencumbered experimentation of new graduates with the architectural governance and quality-control mastery of senior veterans to build resilient, 10x engineering organizations.

Did you find this story helpful?

Share it with your friends and colleagues on social media.

Share

Leave a Comment

Your email address will not be published. Required fields are marked *