Output Control in 2025: The Secret to Better Results




The AI Workflow Maturity Framework: Why Half These Projects Fail, and How to Tell Which Half You’re In
Gartner now says at least half of generative-AI projects get abandoned after the pilot stage. MIT puts real-world enterprise success closer to one in twenty. Neither number means the technology doesn’t work — it means most teams are applying it a stage too early. This is the four-stage framework for finding out which stage you’re actually in.
Most teams don’t fail at AI-powered workflow management because the tools are weak. They fail because a Stage 3 tool gets bolted onto a Stage 1 process. The diagnostic below places your team in under five minutes and points to the one thing worth fixing next.
The honest caveat: a lot of the productivity numbers that circulate about AI at work come from vendor-run surveys with an obvious incentive to look good. The independent numbers — Gartner’s own project-failure tracking, an MIT research initiative, a National Bureau of Economic Research working paper — are more sobering. They’re also more useful, because they show what actually separates the small minority getting real value from the large majority that isn’t.
AI-powered output (workflow) management is the use of AI inside the systems a team already relies on to plan, track, and improve its work — not a separate AI product, but a layer of pattern recognition, anomaly detection, and increasingly autonomous action sitting on top of project trackers, documentation, and reporting. It only produces value once the underlying process is documented and the underlying data is trustworthy. Maturity, not tool selection, is the actual bottleneck for almost every team that tries this and gets disappointed.
“Output management” gets used for everything from print-queue software to team productivity dashboards to AI writing assistants. For the search most people run when they land here, it means one specific thing: the systems a team uses to plan, track, measure, and improve the work it produces. That’s been true for a decade. What changed in the past two years is the AI layer now sitting inside those systems — not as a gimmick, but as something that genuinely changes how a manager sees a bottleneck and how a team self-corrects.
The change is real. It’s just not evenly distributed, and the gap between the teams getting value and the teams getting a bigger invoice is now large enough that three separate research organizations — Gartner, MIT, and the National Bureau of Economic Research — have independently tried to measure it in 2025 and 2026. Their numbers don’t agree on the decimal point, but they agree on the shape: a small share of organizations are seeing real returns, and everyone else is roughly where they started, just with a new subscription.
That gap is the subject of this guide.
That last pairing is the whole argument in miniature. The organizations getting real value aren’t the ones with the biggest AI budget. They’re the ones who changed how work moves before they added the model.
The Four-Stage Maturity Framework
Before any tool recommendation is worth reading, you need to know where you’re actually standing — not where you’d like to be standing. The framework below maps process maturity to the AI intervention that’s actually appropriate for it. It’s a synthesis built from the failure patterns Gartner, MIT, and McKinsey each documented independently in 2025, not a single firm’s proprietary model, which is part of why the same four stages show up under different names in most of the primary research cited throughout this piece.
| Stage | What this looks like | Appropriate AI layer | Representative tools | Biggest mistake |
|---|---|---|---|---|
| Stage 1 Ad Hoc |
Workflows live in people’s heads. Status updates happen over Slack or in meetings. No standard process for recurring tasks. | None yet. Standardize first. | Notion, Trello, a shared doc | Buying an AI analytics platform before there’s clean data to analyze |
| Stage 2 Documented |
Processes are written down. Tasks live in a project tool. KPIs are defined, even if tracking them is still manual. | Rule-based automation for repetitive handoffs | Asana, ClickUp, Monday.com, Zapier | Automating a broken process — you get the same broken result, faster |
| Stage 3 Measured |
Historical data exists. You know cycle times, error rates, and where things stall. Reports go out regularly, even if someone still has to build them by hand. | Predictive analytics, anomaly detection, AI-generated workflow suggestions | ClickUp Brain, Asana AI Studio, Tableau, Power BI | Trusting an AI-flagged pattern without checking whether the underlying data actually supports it |
| Stage 4 Adaptive |
AI suggestions get acted on routinely. The team iterates on process using data, not just intuition. Leadership works from a live dashboard, not a Friday email. | Agent-assisted workflow adjustment, AI-drafted retrospectives, cross-system insight | Microsoft 365 Copilot, Salesforce Agentforce/Einstein, LinearB | Automating judgment calls that genuinely need a human, and quietly eroding the team’s sense of ownership |
Most teams overestimate their own stage by one. If your processes are written down but the weekly report is still someone pulling numbers into a spreadsheet by hand, you’re Stage 2, not Stage 3 — and that distinction determines which tools are actually worth paying for.
Question 1: Could a new hire understand how your team’s core recurring process works without asking a colleague? If no → Stage 1.
Question 2: Can you pull your team’s average task-completion time for last quarter without opening a spreadsheet and doing math? If no → Stage 2.
Question 3: Have you changed a process in the last 90 days specifically because data suggested it was underperforming? If no → Stage 3 or below.
Answered yes to all three? You’re Stage 4-ready. Skip to the tool recommendations below.
You can’t automate a process you haven’t written down, and an AI tool can’t fix a pattern it’s reading out of dirty data.
What the Research Actually Shows
Three independent bodies of research tried to put a real number on AI’s effect on work in 2025: Gartner tracks enterprise project outcomes, MIT’s Project NANDA ran structured interviews and a multi-industry survey specifically on generative AI pilots, and the National Bureau of Economic Research surveyed company executives directly about employment and productivity effects. None of them were trying to sell anything. Their numbers are more sobering than most of what circulates in vendor marketing — and more useful, because the gap between the failures and the small number of successes is where the real lessons live.
AI-powered workflow tools deliver 20–35% efficiency gains more or less automatically once you turn them on.
McKinsey’s 2025 global survey found only about 6% of organizations qualify as “high performers” seeing more than 5% of EBIT attributable to AI. Figures like 20–35% tend to come from samples of organizations that had already redesigned their workflows — the step most teams haven’t taken yet.
Roughly 30% of AI projects get abandoned — an manageable, if unfortunate, minority.
That 30% figure was Gartner’s July 2024 prediction for what would happen by the end of 2025. Gartner’s own account of what actually happened, published in January 2026, put the real figure at 50% or higher — the abandonment rate got worse, not better, as more low-maturity organizations joined the pilot rush.
If 95% of enterprise GenAI pilots are failing, the technology is mostly hype.
MIT’s own researchers draw the opposite conclusion. The failures cluster around generic, poorly-integrated deployments; the 5% that succeed share specific, learnable traits — a narrow use case, workflow-native integration, and usually a vendor partnership rather than an in-house build. That’s a solvable adoption problem, not a verdict on the underlying models.
A note on where these particular numbers come from, since it matters for how much weight to put on any of them: a lot of what circulates online about AI project-failure rates traces back to that single Gartner prediction from mid-2024 — the 30% figure — repeated as if it were a settled finding rather than a forecast for a date that has since passed. Checking Gartner’s own site for what it reported actually happened, rather than relying on outlets that quoted the original forecast, is what turned up the higher, more recent number. Coverage of the MIT report has a similar wrinkle: outlets don’t agree on the exact sample size behind the 95% figure — some cite roughly 150 leadership interviews, MIT’s own published methodology says 52 organizational interviews plus 153 senior-leader survey responses gathered at industry conferences. That kind of drift is a reasonable sign that a given piece of coverage is relaying something secondhand rather than reading the primary report. Where sources disagreed on a number without it changing the substance of the finding, this piece uses the primary source or states the range rather than picking whichever figure sounded cleanest.
One more data point worth sitting with, because it comes from an unusually rigorous source: an NBER working paper surveying company executives directly — not vendors, not AI teams with a reason to talk their own project up — found that more than 90% reported no measurable effect of AI on their firm’s employment over the prior three years, and 89% reported no measurable impact on labor productivity. Among the minority who did report an effect, the average productivity gain was 0.29%. That’s not a rejection of the technology; it’s a picture of how little of it has actually been absorbed into how work gets done, at most organizations, so far. The teams beating that average are, again, the ones that changed the process, not just the software.
What AI is actually good at here — and what it isn’t
AI tools are genuinely strong at three things inside workflow management: finding patterns in historical data faster than a person scanning a spreadsheet, surfacing anomalies that would otherwise hide in the noise of a busy week, and automating the notifications and handoffs that let things fall through cracks in the first place.
They’re not good at knowing whether a delayed task is delayed because the process is broken or because the person doing it just had a rough week. That distinction matters, and it’s still a human manager’s job to make the call. The live question for most teams isn’t whether to use AI in the workflow — it’s whether the people running the team have the standing and the judgment to override it when overriding it is the right call.
Zillow Offers: what happens when the algorithm outruns the process
The clearest illustration of “AI deployed ahead of process maturity” isn’t a small team’s project-tracker rollout — it’s Zillow’s 2021 shutdown of Zillow Offers, its algorithm-driven home-buying arm. The mechanism is the same one that trips up much smaller teams; the scale just makes it impossible to miss.
Zillow’s iBuying model used its Zestimate pricing algorithm to make automated offers on homes, renovate them, and resell at a margin. Through 2021’s overheated, fast-moving housing market, the algorithm’s price forecasts — built for stable, slower-moving conditions — couldn’t keep pace with how quickly regional markets were shifting, especially in hot markets like Phoenix and Atlanta. Zillow ended up holding thousands of homes it had paid above-market prices for.
CEO Rich Barton’s own explanation, from the company’s November 2021 announcement, was direct: the unpredictability in forecasting home prices far exceeds what we anticipated. Stanford Graduate School of Business researchers who studied the collapse afterward found the deeper issue wasn’t the model’s math — it was that Zillow was, by its own later account, a late entrant competing for standardized “cookie-cutter” homes, and started bidding above what its own algorithm recommended to win inventory, which is a process and governance failure sitting on top of an algorithmic one.
The lesson isn’t that predictive algorithms don’t work. It’s that an algorithm making automated, capital-committing decisions needs a process mature enough to catch it the moment its assumptions stop holding — and at Zillow’s scale, by the time humans caught it, the exposure was already in the hundreds of millions.
Most teams reading this aren’t buying thousands of houses on an algorithm’s say-so. But Gartner’s own breakdown of why GenAI projects fail lists the same root causes at any scale: data that isn’t ready, costs that escalate faster than anyone tracked them, and a lack of change management that means even a technically sound tool goes unused. The dollar amounts shrink; the failure mechanism doesn’t.
The Tactical Playbook, by Stage
This is the part that’s supposed to be actionable — “use AI-powered project management tools” isn’t advice, it’s a category name. What follows is organized by the stage transition it’s meant to support, with pricing as published at the time of writing; SaaS pricing changes often enough that it’s worth confirming current numbers directly before you buy.
Stage 1 → 2: Get the process out of people’s heads
You don’t need AI yet. You need documentation and visibility.
Stage 2 → 3: Add real measurement before adding AI
Once tasks are tracked consistently, you can start pulling honest data. This is the step most teams skip — jumping to an AI dashboard before there’s 60–90 days of clean history behind it — which is a large part of why Gartner’s own failure-point research lists “data isn’t ready” as one of the top reasons GenAI projects get abandoned.
Stage 3 → 4: Predictive and agentic AI, now that there’s something to point it at
At Stage 4, the historical data is clean, KPIs are tracked without a manual pull, and managers already have a habit of acting on what the data says. That’s the point at which AI can meaningfully anticipate what’s coming, rather than just narrating what already happened.
Microsoft’s own Copilot pricing changed at least twice in the twelve months before this was published — the enterprise add-on has held near $30/user/month, while the SMB tier dropped from $30 to $21, then to an $18 promotional rate. Treat every price in this section as a starting point for a conversation with the vendor, not a number to build a budget around without checking first.
The Three Things That Actually Move ROI
Across the primary research behind this guide, the same three factors show up in every account of a deployment that actually worked.
- Clean input data before any AI tooling. The Zillow case above is the pattern at extreme scale; Gartner’s own list of GenAI failure points puts “data isn’t ready” at or near the top for a reason. AI tools surface whatever pattern is in the data they’re given — including a pattern that’s actually just inconsistent timestamps or mismatched status labels. Standardize those conventions before adding an analytics layer. Budget four to six weeks. It’s unglamorous and it’s the part that actually works.
- Manager involvement, not just tool access. Gallup’s most recent workplace data makes this concrete: employees whose managers actively support their team’s AI use are 8.7 times more likely to say AI has genuinely changed how much work gets done, and 7.4 times more likely to say it’s freed them up for higher-value work. Engagement — already down for two straight years globally, driven mostly by a drop in manager engagement — climbs to 53% specifically in workplaces where frequent AI use, a clear rollout plan, and active manager support all show up together. Handing out licenses without that combination is one of the most consistent patterns behind a stalled rollout.
- A small pilot before a wide rollout. One team, 60 days, a real before-and-after measurement of cycle time and error rate on that team’s primary deliverable. If the impact can’t be measured on one team in 60 days, it won’t be possible to justify the spend across the organization in 12 months, no matter how the vendor’s case study reads.
The AI didn’t fail these teams. Most of them never gave it a clean enough process to succeed on.
Where This Is Headed: 2027 and the Agentic Shift
Three signals from the same body of 2025 research, read together, point somewhere specific.
First: Gartner’s most recent forecast puts real numbers on the next wave. By 2028, at least 15% of day-to-day work decisions are expected to be made autonomously by agentic AI, up from effectively 0% in 2024, and 33% of enterprise software is expected to ship with agentic capability built in, up from under 1% in 2024. The same research is equally direct about the risk: more than 40% of agentic AI projects started now are expected to be canceled by the end of 2027, for the same reasons killing today’s generative-AI pilots — escalating cost, unclear business value, and risk controls that were an afterthought.
Second: McKinsey’s 2025 data shows the gains from current AI tooling are concentrated almost entirely at Stage 3 and above — the minority of organizations that redesigned a workflow rather than layering AI on top of an unchanged one. Most organizations, by every measure in this piece, are still at Stage 1 or 2.
Third: the cost of Stage 3-level tooling keeps falling. Platforms charging $50–75 per user in 2022 now offer comparable functionality with AI features bundled into base tiers costing a third of that.
Put together, these three signals suggest a specific squeeze coming by 2027–2028: the price of AI-capable workflow tooling keeps dropping while the capability keeps expanding, but the organizations positioned to actually benefit won’t be the ones who bought earliest — they’ll be the ones who used this window to build clean data and manager capability. When agentic AI becomes affordable at the SMB level, it will need a process mature enough to hand real decisions to. Teams that skip that groundwork now are signing up for a catch-up problem no tool purchase will fix later.
Your Next Step This Week
Take the five-minute diagnostic above seriously — most teams place themselves a stage higher than they actually are. If it puts you at Stage 2, the single most valuable thing to do in the next 30 days has nothing to do with software: document your three most important recurring workflows in enough detail that someone new could follow them without asking a question.
That’s not glamorous work. It’s also what every AI deployment in this piece that actually worked was built on top of.
One concrete action to start today: pick your team’s single highest-volume recurring task and time it, start to finish, across three real instances next week — actual clock time, not a guess. That number is your baseline. You’ll need it the next time a vendor’s case study claims a specific percentage improvement.
Frequently Asked Questions
What is AI-powered output management, exactly?
Why do most AI workflow projects fail?
How long does it actually take to move from Stage 1 to Stage 3?
Do I need a dedicated analyst before using any AI tools?
Is it worth paying for the AI add-on in tools like ClickUp or Asana right now?
What’s the difference between workflow automation and AI-powered workflow management?
Should my team start using agentic AI yet?
Glossary
- Workflow maturity
- How standardized, documented, and data-backed a team’s recurring processes are — the framework this guide is built around.
- Proof of concept (POC)
- A limited-scope trial of a technology before a full rollout decision. Most GenAI project abandonment, per Gartner, happens right after this stage.
- EBIT impact
- Earnings before interest and taxes attributable to a specific initiative — McKinsey’s bar for calling an organization an AI “high performer” is more than 5% of EBIT.
- Agentic AI
- AI systems that take multi-step action toward a goal with limited human review at each step, as opposed to responding to a single prompt.
- Shadow AI
- Employees using personal AI tools (ChatGPT, etc.) for work tasks outside of any officially sanctioned pilot — common even at companies where the official pilot has stalled.
- GenAI Divide
- MIT Project NANDA’s term for the split between the roughly 5% of enterprise pilots generating measurable value and the roughly 95% that aren’t.
- AI-ready data
- Data that’s curated, consistently labeled, and well-governed enough for an AI system to reason over reliably — the prerequisite most teams skip.
- RAG (retrieval-augmented generation)
- A technique where an AI model pulls from a specific, current knowledge base before generating a response, rather than relying only on its training data.
https://www.bestprompt.art/how-to-use-chatgpt/
From the BestPrompt.Art Community
https://www.bestprompt.art/how-to-write-ai-prompts-for-beginners/
The maturity framework above applies to creative workflows too—image generation pipelines, content calendars, and collaborative art projects all pass through the same stages. These forum threads map the principles to practice:
Collaborative Art Project: Build a Story Through AI Art The Stage 1 → 2 transition — getting processes out of people’s heads — is exactly what this thread documents. Community members building multi-artist narratives learned quickly that “everyone just contributes when inspired” produces gaps, not stories. The teams that succeeded built explicit handoff conventions first.
https://www.bestprompt.art/what-is-an-ai-prompt/
Prompt Swap: Share a Prompt and See How Others Interpret It. A live demonstration of why standardized documentation matters. When five people run the same prompt across different models and settings, the variance is instructive—and only trackable if everyone records their parameters. This is the creative-domain equivalent of clean data entry before analytics.
https://www.bestprompt.art/ai-prompts-that-actually-work-for-readers/
Top Tools and Resources for AI Artists. The tool recommendations in this post are workflow-management oriented; this community thread covers the creative tooling layer. The same maturity logic applies—don’t pay for AI tier features until you have 60–90 days of consistent usage data to train them on.
https://www.bestprompt.art/how-to-use-keywords-in-ai/
Daily Prompt Challenge: Create Art Based on Today’s Theme! A lightweight version of the pilot-before-rollout principle. One team, one deliverable, 60 days of measurement — compressed into a daily practice cycle. The feedback loop is faster, but the discipline is identical: establish a baseline, introduce a change, measure the delta.
