Watson for Oncology (IBM) — AI cancer treatment recommender commercialised on thin training substrate
2011–2022 · Catastrophic Failure · scored under OTA methodology v4
Scoring
Attribution weights under OTA methodology v4. Percentages express how much of the episode’s outcome each phase and modality accounts for — not a performance grade.
Phase attribution
Observe Easy-Correct · Think Easy-Wrong · Act Easy-Wrong
Modality weights
Modalities scored at zero weight are omitted; the case narrative records why an evidenced modality carries no independent weight.
- Primary modality
- Culture
- Reliability band
- Moderate
- Fraud-related
- No
1. Episode summary
After Watson's 2011 Jeopardy victory, IBM pivoted the platform into healthcare and positioned oncology decision support as the flagship application, described publicly as a "moon shot". Two programmes defined the episode: a partnership with the University of Texas MD Anderson Cancer Center (2013) to build an Oncology Expert Advisor, and a separate training relationship with Memorial Sloan Kettering Cancer Center feeding the commercial product, Watson for Oncology, sold to hospitals in the United States, South Korea, India, Thailand, Slovakia and elsewhere. Both tracks collapsed. A 2016 University of Texas System internal audit documented roughly US$62 million in MD Anderson spend, procurement irregularities involving PricewaterhouseCoopers consulting fees that rose from under US$2 million to above US$20 million, and a gift-fund deficit of approximately US$11.6 million; MD Anderson shelved the project before any patient use. A 2017 STAT News investigation, followed by 2018 STAT and Boston Globe reporting on internal IBM presentations by former deputy chief health officer Andrew Norden, described Watson for Oncology producing treatment recommendations characterised internally as "unsafe and incorrect", including a proposed bleeding-risk drug for a patient already bleeding severely. The commercial system had been trained by a small number of MSK physicians on synthetic, non-real patient cases for each tumour type. IBM sold Watson Health's data assets to Francisco Partners for approximately US$1 billion in 2022, closing the episode. The strategic question the episode turned on: could a curated-expert training substrate, commercialised at the speed IBM's revenue narrative required, carry the generalisation load a global clinical decision-support product actually demanded?
2. Sources
Primary:
- University of Texas System Audit Office, "Special Review of Procurement Procedures Related to the M.D. Anderson Cancer Center Oncology Expert Advisor Project", November 2016 (internal audit report; documents procurement process, PwC fees, gift-fund deficit).
- Ross, Casey and Swetlitz, Ike, "IBM's Watson supercomputer recommended 'unsafe and incorrect' cancer treatments, internal documents show", STAT News, 25 July 2018 (contemporaneous investigative report quoting internal IBM slide decks by Andrew Norden, June/July 2017 presentations).
- Ross, Casey and Swetlitz, Ike, "IBM pitched its Watson supercomputer as a revolution in cancer care. It's nowhere close", STAT News, 5 September 2017 (multi-hospital field investigation including South Korea, Slovakia, Florida sites).
- IBM Newsroom / Francisco Partners, "Francisco Partners to Acquire IBM's Healthcare Data and Analytics Assets", 21 January 2022 (press release documenting divestiture; Merative rebrand).
- The Boston Globe (Weisman, Robert), "IBM documents raise alarm over Watson's diagnostic shortcomings", 29 July 2018 (independent reporting on the leaked internal IBM presentations).
Secondary (with justification):
- Herper, Matthew (synthesising contemporaneous coverage); Schmidt, Charles, "M. D. Anderson Breaks With IBM Watson, Raising Questions About Artificial Intelligence in Oncology", JNCI: Journal of the National Cancer Institute 109(5), May 2017 (peer-reviewed synthesis of the MD Anderson termination, audit findings, and clinical-AI implications).
- Strickland, Eliza, "IBM Watson, heal thyself: How IBM overpromised and underdelivered on AI health care", IEEE Spectrum, April 2019 (investigative feature synthesising interviews with IBM staff, customers, and AI researchers across the Watson Health portfolio).
- Schmidt, Charles, "M. D. Anderson Breaks With IBM Watson", Medscape Oncology ("Big Data Bust"), February 2017 (aggregates post-audit reporting and clinician commentary).
- Dolfing, Henrico, "Case Study 20: The $4 Billion AI Failure of IBM Watson for Oncology", henricodolfing.com, December 2024 (project-management retrospective synthesising primary reporting and public IBM disclosures).
Tertiary (flagged):
- Harvard Business School teaching case, "IBM Watson at MD Anderson Cancer Center" (case-teaching material; frame only, not load-bearing).
3. OTA narrative
Observe. The signals that Watson for Oncology's training substrate was narrow, and that real-world oncology records resisted ingestion, were available to IBM early and from multiple directions. MD Anderson researchers working on leukaemia records reported that the data were frequently missing, ambiguous, or out of chronological order — an empirical observation surfaced during the project, not a post-hoc reconstruction. MSK's training methodology, relying on a small number of expert-authored synthetic cases rather than population-scale real patient records, was visible internally and later surfaced explicitly in the Norden slide decks. Customer feedback from deployment sites — including the reported Jupiter Hospital physician characterisation of the product — reached IBM executives before the STAT reporting. The observation task was routine for the peer group of large clinical-AI programmes of the period: comparable efforts had already flagged record-quality and training-breadth issues as central. Observe was not a root cause in the strategic failure; the signals were in hand. Where Observe fell short it did so as a transmission step — the apparatus produced the signal, and the signal was passed through to the decision layer rather than absorbed there.
Think. The reasoning layer was the primary root-cause phase. IBM and its partner institutions interpreted "trained by world-class oncologists on curated cases" as functionally equivalent to "validated decision support for heterogeneous global patient populations", and committed to commercial scaling on that premise. The correct frame — that synthetic-case training on a narrow expert panel generalises poorly to out-of-distribution real-world records, and that clinical decision support requires prospective outcome validation before marketing claims — was available in the contemporaneous clinical-AI and evidence-based-medicine literature the partner institutions themselves participated in. It was not applied to the commercialisation pathway. The reasoning failure was therefore an Easy-Wrong Think: the interpretive framework existed, was accessible to a reasonably-resourced peer, and was not brought to bear on the go-to-market decision. A secondary reasoning error compounded it — treating a consumer-facing brand association (Jeopardy, "moon shot") as a substitute for clinical validation evidence. Think is a root-cause phase in this episode.
Act. Execution was also a root-cause phase, at the easy end of the difficulty axis. The MD Anderson procurement record documented by the University of Texas audit — single-source awards without competitive process on six of seven service agreements, PwC fees escalating from roughly US$2 million to over US$20 million, an US$11.6 million gift-fund deficit — describes routine governance checks that were omitted, not capability the organisation lacked. Commercial deployment compounded the execution failure: hospitals were sold and marketed the product without IBM publishing peer-reviewed evidence of clinical effect, and the synthetic-case training substrate was not redesigned before global rollout despite internal knowledge of its limits. Act is a root-cause phase, classified at the easy end: the omitted moves — competitive procurement, published validation studies, staged deployment with outcome tracking — were standard for the archetype's peer group. This is therefore a two-phase failure (Think + Act) with Observe carrying the signal through as a transmission step.
4. Modality evidence
Direction. The directional bet on Watson for Oncology was specific, dated, and attributable at the most senior level. The Jeopardy! victory in February 2011 served as the public pivot point; within the same year IBM began repositioning Watson as a healthcare platform, with cancer identified by 2013 as the flagship application. CEO Ginni Rometty declared on a 2015 Charlie Rose appearance that healthcare was "our moon shot," naming the programme publicly as IBM's primary strategic vehicle for escaping years of declining quarterly revenue (Fortune, "Here's How IBM Watson Health Is Transforming the Health Care Industry"; Strickland, IEEE Spectrum, April 2019). The partnership with Memorial Sloan Kettering, formalised in 2012 and expanded into the commercial Watson for Oncology product, was a top-level strategic commitment rather than a skunkworks experiment: both Rometty and former IBM chairman Lou Gerstner held seats on MSK's board of trustees during the development period, creating organisational alignment between the two institutions at the executive level and making the training-partnership architecture a product of concurrent strategic will on both sides (Fortune, ibid.). The decision to commercialise globally before peer-reviewed clinical validation existed — the go-to-market posture documented in the STAT and Boston Globe reporting — is the downstream expression of that directional bet, not a late execution deviation from it.
The direction evidence passes the Step 1 admissibility test: the Rometty "moon shot" declaration is specific (public Charlie Rose interview), datable (2015), and attributed to an identifiable decision-maker cited in a reputable secondary source. The MD Anderson partnership was also dated and attributable: the June 2012 contract is documented in the UT System Audit. Direction is therefore admitted into the modality set, though the weight it carries relative to other modalities is an independent comparative judgement for raters.
Structure. The structural arrangement of the Watson for Oncology programme had two features that are causally relevant to the failure. First, the responsibility for clinical validity of treatment recommendations was placed, in practice, with MSK's training physicians — a small panel of oncologists producing synthetic cases for each tumour type — without an independent validation body or regulatory checkpoint sitting between that panel and global commercial deployment (STAT, Ross and Swetlitz, 25 July 2018; Norden slide decks as reported by STAT and Boston Globe). There was no structural mechanism that would have required IBM to publish peer-reviewed outcome evidence before marketing the product to hospitals in South Korea, India, Thailand, and Slovakia. The commercial product reached markets through a sales and channel structure that treated training provenance (MSK-trained) as equivalent to clinical validation, a design choice visible in IBM's own marketing materials.
Second, IBM Watson Health was formalised as a separate business unit in April 2015 with Deborah DiSanzo, former CEO of Philips Healthcare, as its first general manager, headquartered in Cambridge, Massachusetts, with more than 2,000 employees (Becker's Hospital Review; Strickland, IEEE Spectrum, April 2019). This structural formalisation occurred after the commercial product was already being deployed globally, meaning the governance layer responsible for clinical product accountability was constituted after the commercial trajectory was already set. At MD Anderson, the structural failure was more acute: the UT System Audit documents that the programme was overseen primarily by Dr. Lydia Chin as programme director without adequate institutional oversight, and that the original six-month, $2.4 million contract was extended twelve times without competitive rebidding, with PricewaterhouseCoopers accumulating approximately $23 million in consulting and project-support fees under non-competitive awards (UT System Audit, November 2016; Dolfing, December 2024).
Scoring note (zero-modality rationale): the structural arrangements described in this subsection are classified primarily under Culture in the scoring record on the rationale that the strategic failure causation derived from the behavioural defaults that shaped how the formal architecture was used rather than from a novel divisional architecture or governance design (Watson for Oncology retained a conventional reporting hierarchy across the episode). The dedicated structural elements are counted as the operational substrate of the Culture modality rather than as an independent Structure contribution. Categorisation under METHODOLOGY-ota-scoring-v4.md §5: classification boundary with an adjacent modality. This follows the S-006 (Cisco) precedent for Structure-as-Processes-substrate.
Processes. The validation processes that should have connected training to clinical deployment were either absent or bypassed at every material stage. At the training level, the MSK methodology relied on synthetic cases — hypothetical patients devised by a small number of MSK oncologists — rather than real patient records, and the number of cases used for each tumour type was, as Norden's June and July 2017 internal presentations explicitly documented, "determined without statistical input" (STAT, Ross and Swetlitz, 25 July 2018). The process for quality-controlling recommendations before deployment did not include a mechanism to check whether MSK's treatment preferences diverged from standard clinical guidelines in other healthcare systems — a gap that became acutely visible when international deployment sites found Watson for Oncology recommending treatments reflecting MSK's institutional preferences, which differed from local standards (STAT, Ross and Swetlitz, 5 September 2017, multi-hospital field investigation).
The feedback loop between deployment-site experience and product update was also broken. Customer reports — including the Jupiter Hospital, Florida physician's direct communication to IBM executives that the product was "worthless" and had been purchased "for marketing" — were reaching IBM before the STAT reporting in 2017 and 2018, yet did not trigger a product redesign or a halt to new customer acquisition (STAT, Ross and Swetlitz, 25 July 2018; Strickland, IEEE Spectrum, April 2019). At MD Anderson, the absence of a re-competition process for expanding consulting scope — the contract extended twelve times — reflects a procurement-process failure rather than a capability gap: the institutional procedures for competitive bidding existed in university procurement policy and were simply not applied (UT System Audit, November 2016).
Capability. IBM's NLP capability was genuine but was applied to a task that exceeded its validated scope. Watson's natural language processing could read and synthesise structured clinical literature and extract concepts from physician notes in constrained conditions — capabilities demonstrated in the Jeopardy! domain and in structured question-answering tasks. What the MD Anderson leukaemia project exposed was that real clinical records were frequently missing, ambiguous, or out of chronological order in ways that required domain expertise to interpret, and that the NLP pipeline could not reliably ingest them at the level of accuracy clinical decision-support requires (Strickland, IEEE Spectrum, April 2019; Schmidt, JNCI, May 2017). The capability gap was therefore not in NLP per se but in the organisation's ability to assess the boundary of that capability accurately — to determine in advance what out-of-distribution record structures and international clinical note formats would do to system performance. That meta-capability — knowing what a model cannot yet do — was absent.
IBM did not lack oncological expertise among its partners: MSK physicians and MD Anderson researchers were substantively engaged. The gap was in the institutional capacity to design and execute prospective clinical validation studies that would have made the capability boundary visible before commercial deployment. The Norden slide decks document that IBM's own deputy chief health officer reached this conclusion internally by mid-2017; the slide content on training methodology flaws and the absence of statistical input into case selection reflects a capability audit that arrived too late to affect the go-to-market trajectory (STAT, Ross and Swetlitz, 25 July 2018; Boston Globe, Weisman, 29 July 2018).
Scoring note (zero-modality rationale): the capability described in this subsection is recorded at zero per cent in the modality weights on the rationale of insufficient causal weight — the §4 evidence establishes that Watson for Oncology possessed the technical and operational capability the situation required; the failure mechanism was located in Direction, Processes, Culture rather than in a capability gap. The capability is acknowledged as present in the narrative but does not carry standalone weight in the failure attribution. Categorisation under METHODOLOGY-ota-scoring-v4.md §5 "Zero-modality rationale rule": insufficient causal weight.
Culture. The cultural pathway the primary and secondary sources describe is one in which revenue pressure and brand narrative consistently overrode internal signals of inadequacy. IBM's cognitive solutions division failed to end a streak of twenty-one consecutive quarters of declining revenue, and Watson's healthcare pivot was positioned internally and externally as the means of reversing that trajectory (Strickland, IEEE Spectrum, April 2019; Slate, January 2022). This created a normative environment in which the marketing of Watson for Oncology ran, as contemporaneous accounts characterise it, "way ahead of the capabilities" — not as an isolated communication error but as a persistent organisational pattern (Strickland, IEEE Spectrum, April 2019). The STAT investigations document that IBM executives were receiving clinical-failure signals from deployment sites before the 2017 and 2018 reporting, yet commercial expansion continued. This is the Processes/Culture discriminating question applied: the formal mechanisms for receiving customer feedback existed; the cultural environment determined whether those signals were treated as product-redesign triggers or managed as relationship issues to be contained.
The Norden episode is the clearest cultural marker. Andrew Norden, IBM Watson Health's deputy chief health officer, prepared internal slide presentations in June and July 2017 describing Watson for Oncology's recommendations as reflecting "serious questions about the process for building content and the underlying technology," and documenting specific dangerous recommendation examples — including a drug with a "black box" bleeding warning recommended for a patient with severe bleeding (STAT, Ross and Swetlitz, 25 July 2018). These presentations were shared widely with Watson Health management. The fact that they were produced internally, shared with senior management, and did not produce a product halt before the STAT reporting became public is consistent with a cultural norm in which safety-signal disclosure existed as a documented act but did not carry sufficient organisational weight to interrupt a commercial trajectory set at the CEO level.