This Week in AI Docs: I Graded the Documentation Your Procurement Team Actually Needs. Two Giants Landed on the Same Grade.
Issue 23 | July 1, 2026 | Newsletter | Monday I promised to grade two companies against the five vendor questions. Both hit the exact same score for the exact same reason.
TL;DR: Microsoft 365 Copilot and OpenAI’s ChatGPT Enterprise both land at D- when graded against Monday’s five procurement questions, and it is not because either company hides its subprocessors. Both name them. Both go completely silent on what their AI does poorly and who is responsible for fixing it when it does, and that is the exact gap a procurement team cannot see until the contract is already signed. Grades use the ADECP framework (AI Documentation Ethics and Capability Positioning), scoring Ethical Stakes, Accessibility, Positioning, and Execution. Twelve-minute read.
Monday’s piece walked through five questions to ask any AI vendor before signing a contract with them: where does the data go, what does the AI do poorly, how will you know if it changes, what happens when it makes a mistake, and can the vendor prove regulatory compliance. I said one of the two companies I would grade this week has a nearly universal enterprise footprint, and that the grade surprised me. It did, though not in the direction most people would guess.
THE AI DOCS WATCHLIST
🔴 Microsoft 365 Copilot, Enterprise Subprocessor and Change Documentation (D-, 58%)
What Users Need to Know:
Microsoft 365 Copilot is embedded in the productivity suite nearly every large enterprise already runs, which makes its documentation the closest thing to a universal standard that this framework has ever graded. On the question of where your data goes, Microsoft answers clearly. It publishes a Microsoft Online Services Subprocessors List that names its subprocessors, the third-party companies a vendor quietly routes your data through to actually do the AI processing, states that prompts and responses in Copilot are not used to train foundation models, and confirms that customer content stays inside the Microsoft 365 service boundary. When Anthropic was added as a subprocessor for Copilot experiences, Microsoft documented the exact date it happened, the regions where it defaulted on versus off, and a follow-up admin center update months later when the rollout expanded further. That level of detail is rare. Most vendors, per the DataGrail research Monday’s piece cited, do not disclose their subprocessors at all.
That same rigor disappears entirely on the two questions that actually predict operational risk. Nowhere in Microsoft’s Copilot documentation is there a plain-language list of specific tasks Copilot handles unreliably, the kind of concrete failure scenario Monday’s piece described as the bar a real vendor should clear. Nowhere is there a defined process for what happens when Copilot generates something wrong inside a document, a spreadsheet, or an email a procurement team is relying on, who a customer contacts, or what timeline governs a fix. The documentation that exists is extensive. It answers the questions a legal team asks and stays quiet on the questions an operations team needs answered before rollout.
Where This Information Lives:
Microsoft Online Services Subprocessors List: Names AI subprocessors, including Anthropic, with dated onboarding history and regional defaults.
Data, Privacy, and Security for Microsoft 365 Copilot (Microsoft Learn): States data boundary commitments and confirms Copilot content is not used for foundation model training.
Microsoft 365 admin center, AI providers settings: Gives tenant administrators toggle-level control and default states by region, with dated change history in the Message Center.
No dedicated page addressing specific Copilot failure modes or a customer-facing error and fix process was found during this audit.
Visual Evidence:
A tenant administrator opening the Microsoft 365 admin center under Copilot settings sees a list of AI providers with toggle controls, default states specific to their region, and links to the underlying subprocessor documentation. Nothing in that interface, or in the linked documentation, surfaces a comparable list of known Copilot limitations or a stated path for reporting when the tool gets something wrong inside a live document.
Total: 5/10 points - Partial Disclosure
ADECP Category Breakdown:
Ethical Stakes (25%): 68%. Copilot operates inside email, financial spreadsheets, and internal documents across nearly every enterprise function, which is a high-stakes surface area, and Microsoft’s real data protection commitments meet part of that bar without naming which specific high-stakes tasks Copilot should not be trusted with.
Accessibility (25%): 60%. The subprocessor and toggle documentation is written for IT administrators, not for the broader procurement or legal audience that also needs to evaluate the product, and terms like EU Data Boundary and tenant subprocessor scoping go unexplained for a non-technical reader.
Positioning (20%): 45%. Marketing language about productivity and enterprise readiness is not matched by limitation language given comparable prominence anywhere in the documentation reviewed.
Execution (30%): 55%. The dated, granular change history for subprocessor toggles is a genuine strength. It is undercut by the complete absence of failure mode documentation and error resolution process, which is the highest-weighted category in this framework for a reason.
Weighted Final Grade: D- (58%) Calculation: (68 × 0.25) + (60 × 0.25) + (45 × 0.20) + (55 × 0.30) = 17 + 15 + 9 + 16.5 = 57.5 → D-
The Documentation Reality:
The surprise here is not that Microsoft’s documentation is thin. It is that Microsoft’s documentation is unusually good in exactly the places everyone assumes a giant company would hide the ball, and unusually silent in the place a procurement team actually needs an answer before rollout. Naming Anthropic as a subprocessor, dating the rollout, and giving administrators granular toggle control is real transparency work, and it deserves to be named as such rather than lumped in with companies that bury this information entirely.
That same company has not published a comparably specific answer to what its own flagship product does poorly, or what happens procedurally when it does. Monday’s piece named this exact pattern: a vendor who can describe an error process in a sales meeting but cannot hand over a document that says the same thing does not have a process, they have a talking point. Nothing in Microsoft’s public Copilot documentation contradicts that description, and a footprint this large means the absence affects more procurement teams than almost any other AI vendor decision made this year.
The regulatory compliance signal is real but incomplete. Microsoft references GDPR alignment, Enterprise Data Protection, and the Customer Copyright Commitment, which are substantive artifacts a compliance team can point to. What is missing is anything specific to the EU AI Act’s August 2026 transparency obligations, the exact deadline this newsletter has been tracking for months, applied directly to Copilot’s own documentation.
Limitation Statement Analysis:
Found in Microsoft’s Data, Privacy, and Security documentation for Copilot: “aren’t used to train foundation LLMs.”
Assessment: This is a strong data-use commitment, not a limitation statement, and the distinction matters. It meets the Specific and Action-Oriented parts of the formula for the narrow question of model training. It does not address performance limitations at all, because no comparable statement about what Copilot gets wrong exists anywhere in the documentation reviewed for this audit.
Why D- Instead of D:
A D would suggest a slightly more complete picture than what exists here. Execution carries the most weight in this framework, and Execution is where Microsoft’s documentation fails most completely, with zero evidence of a customer-facing failure mode list or error resolution process for a product used across nearly every function of a large enterprise. The strength on subprocessor disclosure and change tracking keeps this out of F territory, but it cannot offset a total silence on the two questions that matter most once the contract is signed.
What Would Raise the Grade:
Publish a plain-language, regularly updated list of specific tasks Copilot handles unreliably, with real examples, the same rigor currently applied to the subprocessor list. State a defined process for reporting an incorrect Copilot output, including who owns the response and what timeline governs resolution, in a document a procurement team can review before signing rather than a promise made in a sales call. Add an EU AI Act-specific compliance statement addressing the August 2026 transparency obligations directly, rather than relying on adjacent GDPR and data protection language to imply coverage.
Sources: Microsoft Learn, “Overview of AI Subprocessors in Microsoft 365 Copilot,” updated within the past month; Microsoft Learn, “Anthropic models in Microsoft Online Services,” updated within the past month; Microsoft Learn, “Data, Privacy, and Security for Microsoft 365 Copilot,” updated within the past month.
Continue reading below for OpenAI’s ChatGPT Enterprise grade, The Signal on why two unrelated companies landed on the same score, and this week’s full Vendor Documentation Test.





