Who Runs the Project Now?
An AI trained on 750,000 construction projects found $30 million in hidden risk that seasoned human planners missed entirely — but a landmark MIT meta-analysis shows adding a human to that same AI's decision-making can actually make outcomes worse, not better. This episode unpacks the uncomfortable science of who should really run the project: the machine, the manager, or some carefully governed mix of both.
- 00:00 AI's $30 million discovery on the railway
- 02:30 The real question isn't replacement, it's governance
- 04:15 Authority is redistributed, not surrendered
- 06:45 The permissions console is where power actually lives
- 08:30 AI dominates forecasting on historical data
- 10:15 Understanding F1 scores and real AI performance
- 12:30 Humans catch unprecedented shocks AI can't predict
- 14:00 The shocking MIT finding: human-AI combos perform worse
- 16:15 Why adding humans makes forecasting worse
- 18:45 The hybrid synergy debate remains unresolved
- 20:30 Prompt injection is the real security threat
- 22:45 How white text can compromise your project
- 24:30 Building a concrete delegation matrix
- 26:45 Shadow schedulers and the trust problem
- 28:45 The responsibility gap: damned either way
- 31:00 Protect yourself with a decision log
- 33:15 Three key takeaways for tomorrow morning
Read transcript
The Machine That Read 750,000 Projects
Imagine handing your project's future to a machine that has quietly studied 750,000 other people's projects — including all their failures. On Network Rail's Great Western Main Line, that's roughly what happened. An AI model from nPlan, trained on more than 750,000 historical construction schedules representing over $2 trillion of construction spend, was set loose on the project. It hit 75% activity forecasting accuracy and surfaced hidden risks estimated at £30 million in value — risks the human planners, working from experience and gut instinct, hadn't flagged (nPlan case study PDF — Network Rail Great…). When the milestone in question actually completed in May 2019, the outcome fell squarely inside the AI's predicted range (nPlan case study PDF — Network Rail Great…).
Sit with the unease of that for a second. A machine trained on strangers' mistakes could see your project's future more clearly than the seasoned planners standing on site. That's the vivid, slightly unsettling hook at the center of this whole story. And it's not a one-off vendor fairytale. On the HS2 London Tunnels project, ALICE Technologies' AI-driven scheduling optimization is credited with trimming 86 working days and generating £2 million in overhead savings (ALICE Technologies case study PDF — SCS JV…). Peer-reviewed work backs the pattern too: machine-learning models built to predict construction delays achieved strong F1 and accuracy scores that outperformed human expert assessment on large, complex datasets (Applied AI study on predicting constructio…).
Now zoom out. Both the nPlan and ALICE numbers are vendor-published case studies — useful, numeric, but not independently corroborated (nPlan case study PDF — Network Rail Great…)(ALICE Technologies case study PDF — SCS JV…). So treat the exact figures as directional rather than gospel. But the peer-reviewed delay-prediction study lets us anchor the underlying claim on firmer ground: for structured, data-intensive forecasting, AI genuinely does beat unaided human judgment (Applied AI study on predicting constructio…). Crucially, the same study notes the limit — human experts remain better at recognizing unprecedented situations and external shocks, the stuff no training set has seen (Applied AI study on predicting constructio…). That boundary line — pattern-detection versus the genuinely novel — is the fault line this entire episode runs along.
The question this episode asks isn't the tired one, 'will AI take my job?' It's the sharper, more useful one: which of your tasks should you hand over, under what governance, and who is accountable when it goes wrong?
It hit 75% activity forecasting accuracy and surfaced hidden risks estimated at £30 million in value that human planners hadn't flagged.
AI's advantage is task-specific: strong on structured pattern-detection, weak where the training data runs out. Vendor figures (nPlan, ALICE) are directional, not independently verified.
What this means for listeners: If your work involves structured forecasting on data you have history for — schedule risk, delay probability, cost overrun — the honest answer is that a well-trained model will likely beat your gut, and pretending otherwise costs you. The skill worth building isn't out-forecasting the machine; it's knowing exactly where its training data runs out.
24/7 Autonomy, With an Asterisk
Read the marketing pages for today's agentic project tools and you'd think the robots already have the keys. monday.com markets its AI agents with bold '24/7 autonomy' language (monday.com Support — 'AI Agents on monday…). ClickUp names its feature 'Super Agents' and leans hard into autonomy framing (ClickUp product pages — 'Autonomous Agents…). Asana's investor releases describe an 'Operating System for Human-Agent Teams' (Asana Investor Relations PR (Aug 25 2025;…). It sounds like the human is being edged out of the room.
Then read the help-center fine print, and a much more modest picture emerges. monday.com's own support documentation describes an 'AI Permissions and Governance' admin area that categorizes agent types — user agents, monday agents, third-party agents, org agents, external connectors — and lets admins set permissions per feature and per agent type; board-scoped access, where an agent can be granted read access to only particular boards, is a first-class concept (monday.com Support — 'AI Agents on monday…). That is not '24/7 autonomy.' That is a carefully permissioned assistant whose leash length the admin sets. Microsoft's Planner Agent, which reached general availability in mid-to-late June 2026, can create and update tasks via natural language — but with tenant admin consent required for cross-region data movement, and with documented admin controls to turn the agent off entirely (Microsoft Tech Community — 'Planner Agent…). Asana is the most explicit of all, publicly framing its strategy as collaboration over autonomy, positioning its agents as 'governed teammates' with human control, admin visibility, and usage limits rather than hands-off autonomy (Asana Investor Relations PR (Aug 25 2025;…). ClickUp's autonomy-forward pages, on closer read, recommend narrow permissions and comprehensive audit logging, and its help docs describe role gating so that limited roles can't even see or trigger Super Agents (ClickUp product pages — 'Autonomous Agents…).
So what's really going on? The gap between the marketing pitch and the help-center reality is its own small story about how 'agentic AI' is being sold versus how it's actually built. The technical capabilities across vendors are broadly similar — agents that can propose and execute changes within scope, gated by admin-configurable permissions (monday.com Support — 'AI Agents on monday…)(Asana Investor Relations PR (Aug 25 2025;…)(ClickUp product pages — 'Autonomous Agents…). The difference is positioning: Asana chooses to sell control, while monday.com and ClickUp choose to sell autonomy (Asana Investor Relations PR (Aug 25 2025;…)(monday.com Support — 'AI Agents on monday…)(ClickUp product pages — 'Autonomous Agents…). That reflects market strategy, not a difference in what the software can do.
The practical lesson is to calibrate your skepticism accordingly. When a vendor says 'autonomous,' the right follow-up question isn't 'how smart is it?' but 'what exactly can it write to, and who set that boundary?'
That is not 24/7 autonomy — that is a carefully permissioned assistant whose leash length the admin sets.
What this means for listeners: Before you believe any autonomy claim, go find the permissions documentation, not the landing page. The real capability of an agentic tool lives in its admin console — what it can read, what it can write, and what requires a human to sign off — and that is entirely within your control to configure.
The -0.23 Problem: When Teaming Up Makes You Worse
Here's a finding that should make anyone who's ever said 'human plus AI is the best of both worlds' pause. MIT's Center for Collective Intelligence analyzed 370 effect sizes across 106 experiments and found that, on average, human-AI combinations performed significantly worse than the best of humans alone or the best of AI alone (Hedges' g = -0.23, 95% CI [-0.39, -0.07]) (Vaccaro, Almaatouq & Malone (2024). *When…). The dream of 'centaur' collaboration — human and machine fusing their strengths — simply doesn't hold up as a general rule.
If we stopped at that average, it would look like a flat contradiction of the £30 million ghost we met earlier. How can AI-plus-human be worse than AI alone, if hybrid construction forecasting is such a triumph? The resolution is that the meta-analysis didn't stop at the average — it decomposed by task type, and that's where the real insight lives. The researchers found performance losses concentrated in decision-making tasks and significant gains in content-creation tasks (Vaccaro, Almaatouq & Malone (2024). When…). More precisely: when humans outperformed AI alone, the combination produced gains; but when AI outperformed humans alone, the combination produced losses (Vaccaro, Almaatouq & Malone (2024). When…). For decision tasks like classifying deepfakes, forecasting demand, and diagnosing medical cases, human-AI teams often underperformed against AI alone (Vaccaro, Almaatouq & Malone (2024). *When…).
Now the pieces click together. Construction delay and cost-overrun prediction is squarely a forecasting task where AI already has the data advantage (Applied AI study on predicting constructio…). So when a vendor case study describes 'hybrid success,' it is almost certainly not describing the synergistic blending the centaur dream imagines. It's describing AI doing the forecasting work with a human veto stapled on top (Vaccaro, Almaatouq & Malone (2024). When…)(Applied AI study on predicting constructio…). That's the episode's central, genuinely unresolved debate: is 'human-in-the-loop' really collaboration, or is it AI-dominance-with-override, dressed up in collaborative language for comfort and liability cover? This isn't a factual contradiction — it's an interpretive dispute, and no study has yet compared the meta-analytic method head-to-head against a construction-specific dataset using the same effect-size methodology (Vaccaro, Almaatouq & Malone (2024). When…).
There's a sharper edge here worth naming. The 'hybrid outperforms' claim quietly assumes the human override is well-calibrated. But cross-domain evidence from radiology shows that assumption often fails: some clinicians are helped by AI assistance while others are actively hurt, and less-experienced users are not guaranteed to benefit — a poorly-calibrated user may trust a wrong suggestion or get distracted by a right one (Radiology AI-calibration study on differen…). Staple an uncalibrated human onto a strong AI and you don't get the best of both; you get -0.23.
When the AI outperformed humans alone, the combination produced losses — staple an uncalibrated human onto a strong AI and you get minus 0.23.
The rigorous meta-analytic finding (Tier 1) undercuts the vendor 'hybrid synergy' framing (Tier 3). The construction case studies show AI-alone dominance on a forecasting task, not centaur synergy.
What this means for listeners: Don't assume that adding yourself to the loop improves an AI's forecast — on tasks where the AI is already the stronger party, you may be adding noise. Reserve your intervention for the cases where you genuinely know something the training data doesn't, and be honest with yourself about how often that actually is.
The New Attack Surface: When a Trusted Agent Gets Poisoned
Give an AI agent write-access to your Jira, Asana, or Planner board and you've done something genuinely new to your risk profile — but probably not the risk you'd guess. Most procurement checklists worry about classic privilege escalation: an agent breaking out of its permissions to grab access it shouldn't have. The security literature points somewhere subtler and scarier. The dominant risk in agentic deployments is prompt injection and poisoned context causing an already-authorized agent to misuse its legitimate tool access at scale (Help Net Security / OWASP prompt-injection…). The agent never breaks the rules; it's tricked into using the access it was correctly given, to do damage.
This isn't theoretical. RAXE Labs documented a cluster of CrewAI vulnerabilities — CVE-2026-2275, -2285, -2286, and -2287 — chainable from a prompt injection all the way to remote code execution, server-side request forgery, and file reads (RAXE Labs Security Advisory RAXE-2026-049…). The Cloud Security Alliance's March 2026 research note detailed LangChain and LangGraph vulnerabilities under coordinated disclosure, highlighting indirect prompt-injection risk hiding inside tool integrations (Cloud Security Alliance research note — La…). And a December 2025 arXiv preprint ran comparative penetration testing across AutoGen and CrewAI, demonstrating multiple successful attack scenarios: prompt injection, SSRF, SQL injection, and tool misuse (arXiv preprint (Dec 2025) — comparative pe…).
Multi-agent architectures make this worse by multiplying trust boundaries. In an orchestrator-plus-downstream-agent setup, a compromised or manipulated orchestrator can pass flawed instructions to agents that trust it implicitly — and a single poisoned handoff can cascade across the system, altering schedules or misreporting status before any human notices. Every agent you add is a new trust boundary and a new point of failure. (You'll sometimes see the dramatic 'cascading error' framing cited via secondhand blog write-ups of an arXiv paper rather than the paper itself — worth flagging as an unverified pointer, not established fact.)
So the procurement implication flips. The right RFP question isn't only 'can this agent's permissions be broken?' It's 'what happens when this correctly-permissioned agent is fed a malicious instruction through the project data it's allowed to read?' Build your due diligence around agent inventory, risk classification, policy enforcement, observability, and human-oversight configuration (Help Net Security / OWASP prompt-injection…).
The agent never breaks the rules; it's tricked into using the access it was correctly given, to do damage.
What this means for listeners: If you're buying or scaling an agentic PM tool, your security questions need to shift from IAM permissions to context integrity — assume the agent's access is legitimate and ask how it's protected from being manipulated through the data it reads. Demand audit logging, a documented patch SLA for critical CVEs, and a capped autonomy scope until the tool has proven itself.
What PMs Will Hand Over — And What They Won't
Watch how project managers actually behave around AI and a remarkably consistent boundary appears. They readily delegate structured, data-intensive tasks — schedule optimization, resource leveling, risk scoring, standardized status reporting — where there's historical pattern data, low ambiguity, and low reputational or ethical stakes. They cling to the tasks involving strategic trade-offs, stakeholder negotiation, and ethically-charged calls. Multiple sources converge on this split, including PMI's own profession data and construction workflow studies (Applied AI study on predicting constructio…)(Empirical study on fairness perceptions an…).
There's a good reason it's not just habit. Fairness perceptions directly predict behavioral trust in AI decision-making: people accept and rely on AI recommendations they perceive as procedurally fair, and they resist systems they see as biased — even at the cost of reduced accuracy (Empirical study on fairness perceptions an…). Task allocation, performance evaluation, and resource decisions are exactly where fairness concerns bite hardest, which is precisely why PMs guard them. Delegating a fairness-loaded decision to an opaque algorithm feels, to the people affected, like shirking responsibility (Empirical study on fairness perceptions an…).
And here's the ground-truth wrinkle the dashboards don't show. While vendors tout AI-optimized schedules, construction site superintendents reportedly engage in quiet 'shadow practices' — ignoring or working around AI recommendations that don't match their gut sense of what's actually happening on site, as discussed candidly in a Reddit construction thread. That's an appealing 'workers versus the machine' narrative, and it humanizes the governance debate — but be clear-eyed that it's anecdotal, drawn from a single forum thread, not verified data. Still, it points at something real: official adoption narratives and on-the-ground practice can diverge sharply, and if your override data all lives in people's heads, you're flying blind.
The practical move is to make the delegation boundary explicit rather than letting it drift. Classify every project task into a two-tier matrix before deployment. Tier A — delegate to AI: historical pattern data, low ambiguity, low stakes. Tier B — retain human authority: stakeholder negotiation, scope trade-offs, ethical calls. Then instrument it: require human sign-off on any AI-generated change exceeding 5% of timeline or budget, escalate to strategic review when AI confidence drops below 70% where the vendor exposes it, and re-evaluate the classification every 90 days as performance data accumulates.
Require human sign-off on any AI-generated change exceeding 5% of timeline or budget, and escalate when AI confidence drops below 70%.
Classify tasks by how much historical pattern data exists and how high the ethical/reputational stakes are. Delegate the data-rich, low-stakes quadrant; retain the rest.
What this means for listeners: Draw your Tier A / Tier B line on paper before you switch anything on, and put numbers on the escalation thresholds — a 5% timeline-or-budget trigger and a 70% confidence floor are defensible starting points. Then actually log your overrides, because the gap between the official dashboard and what your best people quietly do is where both your risk and your learning are hiding.
Damned If You Do, Damned If You Don't
Here's the accountability trap, borrowed from a field that's lived it longer than project management has. In healthcare, behavioral research is blunt: a physician gets blamed for a bad outcome whether they followed the AI's recommendation or overrode it. Under U.S. malpractice law, courts judge the physician against a 'reasonable physician under similar circumstances' standard, whether or not AI was involved (Legal/behavioral analysis of U.S. malpract…). And behavioral science shows we judge humans much more harshly than AI for wrongness and blameworthiness, an asymmetry expected to intensify rather than fade (Legal/behavioral analysis of U.S. malpract…). Translate that to project management and the bind is exact: override the AI's risk flag and the project fails, you're blamed for ignoring the signal; defer to the AI and it's wrong, you're blamed for failing to exercise judgment.
What makes this a genuine gap rather than just an anxiety is that the formal law is drifting the other way. Liability doctrine is moving toward organizational and vendor responsibility, borrowing the logic of vicarious liability — the way an employer is held responsible for an employee's on-the-job conduct (Legal/behavioral analysis of U.S. malpract…). Regulators have explicitly rejected the 'no one controls it' defense: in the CFTC's action against Ooki DAO, 'no one controls it' didn't shield the token holders from liability, and 'the AI did it' won't shield a deployer either (Legal/behavioral analysis of U.S. malpract…).
Stack those two trends and you get the responsibility gap that will define workplace tension as these tools mature. Formal liability is migrating up toward organizations and vendors. Social and psychological blame stays stubbornly fixed on the individual human operator (Legal/behavioral analysis of U.S. malpract…). The PM is left holding a bag the law says belongs to someone else. And the psychological toll is real even without PM-specific data yet: a general workforce survey found 43% of employees worry they'll be seen as lazy or untrustworthy for relying on AI, and 24% feel judged or second-guessed when using AI help (SnapLogic workplace survey on AI usage anx…). That's the ambient anxiety a PM carries into every override decision.
You can't close the responsibility gap single-handed, but you can refuse to absorb it silently. Write an accountability charter before you deploy: specify in writing which decisions require human sign-off, log every AI-recommended versus human-overridden decision, and define the escalation path when AI and human judgment conflict.
Formal liability migrates up toward organizations, but social blame stays fixed on the individual PM — who is left holding a bag the law says belongs to someone else.
What this means for listeners: Protect yourself with documentation, because in practice the blame lands on you regardless of which way you decide. A written accountability charter and a complete decision log aren't bureaucracy — they're the only evidence that will exist that you exercised judgment rather than abdicated it.
Orchestrator or Rubber Stamp? Building the Role You Actually Want
So who runs the project now? The honest answer is: it's co-run, and the role of the human is being rewritten from planner-controller toward orchestrator, curator, and ethical governor (Annual Review of Organizational Psychology…)(Systematic review of algorithmic managemen…). But there's a real question buried inside that upbeat reframe, and no one has answered it: is 'orchestrator' genuine career elevation, or a euphemized demotion? No direct survey has ever asked project managers whether the term feels like promotion or polish (Annual Review of Organizational Psychology…).
The organizational-behavior literature says the answer isn't fixed — it depends on the architecture of the tool you're handed. Algorithmic management has graduated from a gig-economy edge case into a mainstream organizing principle, with both erosive and augmentative effects depending on organizational design (Annual Review of Organizational Psychology…). The systematic-review evidence sharpens the mechanism: prescriptive, automatic-execution algorithms — where the system maps inputs to outputs and acts, and the human rubber-stamps — drive erosion of managerial authority; advisory, overridable algorithms that managers can disregard, alter, or negotiate with preserve and even enhance discretion by stripping away low-value cognitive load (Systematic review of algorithmic managemen…). Both camps are well-supported. Which architecture actually dominates in real-world PM deployments is genuinely unknown — which is why it's untestable, today, whether erosion or augmentation is winning (Systematic review of algorithmic managemen…).
That turns into a concrete question you can ask yourself right now: does the tool on your screen tell you what to do, or ask you what you think? If it's the former, you're being nudged toward the rubber-stamp end of the spectrum, and the 'orchestrator' language is doing PR work. If it's the latter, the elevation is real. There's a related open question worth watching in your own workplace: is AI becoming a visible, explicitly-governed actor — the direction PMI's new AI Standard for Project Professionals hints at — or is it becoming more embedded and less visible even as it takes on more authority? PMI's own Pulse data suggests AI is often running behind the scenes with human skills still described as central (Applied AI study on predicting constructio…). Honestly, we don't yet know which way that resolves, and it's worth flagging as unresolved rather than pretending otherwise.
Here's how to make the transition responsibly rather than passively. Run a governed pilot before any org-wide rollout, then calibrate trust deliberately, then lock in accountability — in that order.
Does the tool on your screen tell you what to do, or ask you what you think?
Sequence the deployment: prove the tool in a narrow pilot, calibrate human trust, then formalize accountability before scaling. The pilot phase is the one you cannot skip.
What this means for listeners: You have more agency over the orchestrator-versus-rubber-stamp outcome than the marketing implies — it turns largely on whether your tools are advisory or prescriptive, and on whether you insist on override rights. Watch your own workplace for the tell: a tool that asks what you think is elevating you; a tool that tells you what to do is quietly deskilling you.
- Vaccaro, Almaatouq & Malone (2024). *When Combinations of Humans and AI Are Useful: A Meta-Analysis.* Nature Human Behaviour. 370 effect sizes, 106 experiments.
- nPlan case study PDF — Network Rail Great Western Main Line (vendor-published, not independently verified).
- Applied AI study on predicting construction project delays via machine learning (peer-reviewed applied study).
- ALICE Technologies case study PDF — SCS JV, HS2 London Tunnels (single-vendor case study).
- Microsoft Tech Community — 'Planner Agent in Microsoft 365 Copilot is now generally available,' June 2026; Microsoft Learn Copilot/Planner docs.
- monday.com Support — 'AI Agents on monday.com,' 'AI Permissions and Governance,' 'Bring your external agent'; monday.com 'AI Agents' marketing page.
- Asana Investor Relations PR (Aug 25 2025; June 4 2026) — 'Operating System for Human-Agent Teams'; Asana Help Center AI admin controls.
- ClickUp product pages — 'Autonomous Agents / Super Agents'; ClickUp Help — AI feature availability and limits.
- RAXE Labs Security Advisory RAXE-2026-049 — CrewAI vulnerability cluster (CVE-2026-2275/-2285/-2286/-2287).
- Cloud Security Alliance research note — LangChain/LangGraph vulnerabilities, coordinated disclosure March 2026.
- arXiv preprint (Dec 2025) — comparative penetration testing across AutoGen and CrewAI (peer-review status unclear).
- Help Net Security / OWASP prompt-injection summary, June 2026 (secondary synthesis).
- Annual Review of Organizational Psychology and Organizational Behavior — algorithmic management review.
- Systematic review of algorithmic management literature (Edwards et al. 2024; Krzywdzinski et al. 2024; Meijerink et al. 2021a; Neumann et al. 2023).
- Legal/behavioral analysis of U.S. malpractice law and blame asymmetry; CFTC v. Ooki DAO commentary on vicarious liability.
- Empirical study on fairness perceptions and behavioral trust in AI decision-making.
- SnapLogic workplace survey on AI usage anxiety (general employee sample, not PM-specific).
- Radiology AI-calibration study on differential clinician responses to AI assistance (cross-domain analogy).