R8 — Negative Evidence File: The Case Against Building Atlas
Research agent output. Today: 2026-07-18. Scope: prosecution’s evidence only — reasons NOT to build an internal wiki/portal/knowledge-management/agent-memory system (“Atlas”) for a 2-human + 64-agent setup. Every claim below is sourced; weak/uncited secondary sources are flagged explicitly as such. Do not treat flagged items as verified facts — treat them as “commonly repeated claims of uncertain provenance,” which is itself useful prosecution material (shows the KM-failure narrative is folk wisdom as much as data).
1. Wiki / KM / documentation-rot literature
Documentation Rot is a named, recurring pattern, not a one-off complaint.
- dev.to (Kislay, undated, retrieved 2026-07-18), “Why Your Engineering Wiki is a Graveyard (And How to Fix It)”: https://dev.to/kislay/why-your-engineering-wiki-is-a-graveyard-and-how-to-fix-it-2eme — claims: context-switching friction between coding and wiki-updating kills documentation habit; content decay widens silently as “every process update, tool migration, or team reorg makes some percentage of the wiki quietly wrong”; performance systems reward shipping over documenting; “once you encounter one outdated doc, you stop trusting all docs” (trust cascades to zero, not gracefully).
- Pravodha Blog (undated), “Your Wiki Isn’t a Knowledge Base. It’s a Graveyard.”: https://pravodha.com/blogs/your-wiki-isnt-a-knowledge-base-its-a-graveyard — “a Slack explanation written last Tuesday by a senior engineer who knows the system is inherently more credible than a Confluence page last updated before the last two reorgs.” Documentation is structurally a separate activity from the work itself, which is why it always loses the priority fight against a production bug or a Friday deadline.
- The Content Wrangler (undated), “When Confluence Starts Groaning”: https://www.thecontentwrangler.com/p/when-confluence-starts-groaning-a — Confluence groans “not because it’s broken, but because teams have asked it to do a job it was never designed to handle at scale.”
Academic / empirical backing (stronger sourcing):
- ICSE 2019 Technical Track, “Software Documentation Issues Unveiled”: https://2019.icse-conferences.org/details/icse-2019-Technical-Papers/49/Software-Documentation-Issues-Unveiled — mined 878 documentation-related artifacts across mailing lists, Stack Overflow, issue trackers, and PRs; built a taxonomy of documentation problems. Peer-reviewed, real sample size — the strongest single source in this bucket.
- ResearchGate, “Evaluating usage and quality of technical software documentation: An empirical study”: https://www.researchgate.net/publication/262201755 — finds documentation “often outdated, incomplete and sometimes not beneficial.”
- ScienceDirect systematic review (2024), “mapping study on documentation in Continuous Software Development”: https://www.sciencedirect.com/science/article/pii/S095058492100183X — documentation “regarded as a time-consuming and arduous task, which frequently leads to its neglect or obsolescence” across 29 primary studies.
Weakly-sourced but repeated stats (flag: uncited/undated primary sources, treat as folk-numbers, not verified data):
- episteca.ai blog (undated), “The Documentation Decay Problem”: https://episteca.ai/blog/documentation-decay/ — cites (without linkable primary source) “Readme: docs materially outdated within 30-90 days of publication”; “Zoomin: 68% of enterprise technical content untouched >6mo, 34% untouched >1yr”; “Guru: 60% of employees don’t trust their internal KB”; “Salesforce: 78% of support escalations traceable to knowledge gaps”; “IDC: Fortune 500 loses $31.5B/yr to KM failures.” I could not independently locate the original Zoomin/Salesforce/IDC reports behind these numbers in a follow-up search — they may be real but are currently unverifiable via public search. Cite only as “widely-repeated industry claim,” not as fact.
Wiki adoption failure driver — organizational, not technical:
- ResearchGate (undated), “Adoption of Knowledge Management Systems: A Study on How Wiki Systems Should Be Adopted by Minimizing the Risk of Failure”: https://www.researchgate.net/publication/318947821 — across seven case studies of unsuccessful wiki adoption, every single one cited lack of management support as the primary failure driver, not tooling quality. Implication: Atlas succeeding or failing will hinge on whether the 2 humans keep actively championing/feeding it, not on which stack it’s built on.
- Bloomfire, “8 Reasons Why Knowledge Management Fails”: https://bloomfire.com/blog/why-knowledge-management-fails/ — “when knowledge management feels like extra work, employees stop contributing. Systems go stale, adoption plummets, and the initiative quietly fails.”
2. Backstage adoption failure reports
Primary source, freshest and most detailed:
-
Zohar Einy (Port.io co-founder — note: Port is a paid Backstage competitor, so this source has a commercial incentive to bury Backstage; weigh accordingly, but the specific numbers are independently corroborated by the Medium piece below), “Backstage is dead”, Autonomous Engineering newsletter, 2026-03-31: https://newsletter.port.io/p/backstage-is-dead
- Spotify internal adoption: ~99%. External-org average adoption: ~10% (“equivalent to a proof of concept”).
- ShiftMag case study: only 5% of engineers actually used Backstage despite a technically successful deployment.
- Time to a usable instance: 6–12 months. Ongoing cost: 2–5 full-time engineers for years; true TCO estimated at ~$150K per 20 developers.
- Root cause chain: catalog relies on manually-maintained YAML → drifts from reality the moment a service changes → “catalogs go stale fast. Once trust in the catalog erodes, adoption collapses” → org either abandons within 6 months (“install and forget”) or discovers a year+ in “just how hard it is to maintain.”
- Direct quote from an adopter: “The idea of backstage is super cool, but for me, as a DevOps engineer, the fact that I need to write a lot of React code instead of GitHub workflows and Terraform files made me leave the project.”
-
Medium (Samadhi Anuththara, undated), “Backstage Backlash: Why Developer Portals Struggle”: https://medium.com/@samadhi-anuththara/backstage-backlash-why-developer-portals-struggle-cb82d4f082e1 — corroborates ~10% external adoption independently; adds: plugins are “separate widgets rather than integrated experiences,” inconsistent community-plugin quality, portals fail without an “immediate value” proposition and become “solutions in search of a problem” absent strong product ownership.
Verdict pattern across both sources: Backstage failures are consistently described as adoption failures, not technical failures — the tool works; organizations don’t sustain the discipline required to keep it truthful.
3. Agent long-term memory write-back failures (2025–2026)
Security research — memory poisoning as an attack class:
- MINJA (Memory INJection Attack), Dong et al., NeurIPS 2025 poster: https://neurips.cc/virtual/2025/poster/118152 and paper https://arxiv.org/html/2503.03704v4 — attacker injects malicious records into an agent’s memory bank via query-only interaction (no privileged access needed). Reported 95%+ injection success, 98.2% injection success / 76.8% downstream attack success across Webshop, MIMIC-III, MMLU agents. This is peer-reviewed / conference-accepted, the strongest source in this bucket.
- “MemoryGraft” attack, described as published December 2025 (per secondary summary in Medium/InstaTunnel piece — I could not locate the original paper/preprint directly, flag as secondary-sourced): https://medium.com/@instatunnel/agentic-memory-poisoning-how-long-term-ai-context-can-be-weaponized-7c0eb213bd1a — benign-looking content planted in long-term memory, retrieved and imitated by the agent “weeks later” on similar tasks.
- OWASP: memory/context poisoning listed as ASI06, a top agentic-AI risk for 2026 (per WorkOS summary, https://workos.com/blog/ai-agent-memory-poisoning — WorkOS is a vendor, but OWASP list itself is the citable primary claim; I did not independently pull the raw OWASP ASI list text in this pass).
- Survey: “Always-On Agents: A Survey of Persistent Memory, State, and Governance in LLM Agents,” Ding et al., arXiv:2606.30306, dated 2026-06-30: https://arxiv.org/pdf/2606.30306 — 136 pages, 400+ citations, structured around stale-context accumulation, factual-consistency/conflict problems, write-back integrity failures, and governance/access-control gaps as the four core failure families for persistent agent memory. (Full section text not extracted — PDF returned structural metadata only; treat section-level claims as indicative of scope, not verified quotes.)
Real-world/product-level evidence of wrong-fact persistence (non-adversarial — just organic decay):
- techbuzz.ai (undated), “ChatGPT’s Memory Feature Silently Poisons Answers With Bad Data”: https://www.techbuzz.ai/articles/chatgpt-s-memory-feature-silently-poisons-answers-with-bad-data — testing found ChatGPT memory storing outdated assumptions, building “detailed but incorrect personal profiles,” and treating stored wrong facts “with the same confidence it applies to factual knowledge… without flagging uncertainty.”
- OpenAI Developer Community threads (user complaints, dates on threads late 2025/2026): “ChatGPT memory broken at the moment” https://community.openai.com/t/chatgpt-memory-broken-at-the-moment/1108272 ; “[Bug] GPT-4o memory regression — context loss across chats” https://community.openai.com/t/bug-gpt-4o-memory-regression-context-loss-across-chats-and-inside-threads/1310926 — reported: duplicate memory records, “saved memory full” ceiling around 100–150 entries causing silent drops of new (or displacement of old) facts, model unilaterally deciding what’s memory-worthy so users “usually only notice two weeks later when it cannot help them.”
- Coding-agent specific: MindStudio blog (undated), https://www.mindstudio.ai/blog/persistent-memory-system-claude-code-agents — anecdote: three months into using Claude Code on a production codebase, developers were “still correcting the same mistakes every session” (“No, we use pnpm, not npm”) because “every correction vanished when the session ended” and CLAUDE.md-style memory caps out and goes stale — “a memory file written on day 1 will still be in MEMORY.md on day 365 unless the agent or user manually deletes it.” Directly on point for Atlas’s proposed write-back mechanism.
- Academic (empirical, strongest source in this sub-bucket): “Agent READMEs: An Empirical Study of Context Files for Agentic Coding,” 11 authors, arXiv 2511.12884, dated 2025-11-18: https://arxiv.org/pdf/2511.12884 — mined 2,303 real AGENTS.md/CLAUDE.md files across 1,925 repos. Files “evolve like configuration code, maintained through frequent, small additions.” Found functional context (build commands 62.3%, implementation detail 69.9%, architecture 67.7%) heavily emphasized but non-functional/safety content starkly underspecified: security guidance in only 14.5%, performance guidance in only 14.5% of files. Reads as evidence that even actively-maintained agent-memory files systematically omit the categories of guidance most likely to prevent costly mistakes — the files that do get maintained get maintained for convenience, not correctness/safety.
Mechanistic reason large context (which any “Atlas” retrieval layer feeds into) degrades rather than helps beyond a point:
- Chroma Research (2025), “Context Rot: How Increasing Input Tokens Impacts LLM Performance”: https://www.trychroma.com/research/context-rot — tested 18 frontier models (incl. GPT-4.1, Claude Opus 4, Gemini 2.5); every model degrades non-uniformly as input length grows, “sometimes by 30 to 50 percent well before the documented [context] limit.” “Lost-in-the-middle” and distractor interference (semantically similar-but-irrelevant retrieved content actively misleading the model) are named mechanisms. Practical safe-context budget for 2M-token-window models lands at 150K–400K tokens for high-accuracy work — i.e., dumping more of Atlas’s accumulated knowledge into context is not free and has a measured ceiling. https://www.understandingai.org/p/context-rot-the-emerging-challenge is a secondary summary of the same research if the primary Chroma page is paywalled/changed.
4. Duplicate-merge / FAQ auto-answer harms
Real, adjudicated legal precedent for “chatbot gave a wrong answer, company is liable”:
- Moffatt v. Air Canada, 2024 BCCRT 149 (Canadian Civil Resolution Tribunal, decided February 2024). Coverage: Forbes https://www.forbes.com/sites/marisagarcia/2024/02/19/what-air-canada-lost-in-remarkable-lying-ai-chatbot-case/ and McCarthy Tétrault legal analysis https://www.mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot — Air Canada’s website chatbot hallucinated a bereavement-fare policy; tribunal awarded the customer $812.02 and explicitly rejected Air Canada’s argument that the chatbot was “a separate legal entity responsible for its own actions.” Standard of care: a company must “take reasonable care to ensure representations are accurate and not misleading” — full stop, regardless of whether a human or a bot made the representation. This is the load-bearing precedent for “an FAQ-autoanswer being wrong isn’t a neutral bug, it’s an accountability problem for whoever owns Atlas.”
Real, dated, high-visibility internal-support-bot hallucination with concrete business damage (not legal, but reputational + churn):
- Cursor (Anysphere) AI support-bot incident, April 2025. Coverage: The Register https://www.theregister.com/2025/04/18/cursor_ai_support_bot_lies/ , Hacker News thread https://news.ycombinator.com/item?id=43683012 , AI Incident Database entry #1039 https://incidentdatabase.ai/cite/1039/ — Cursor’s support AI (“Sam”) fabricated a “one device per subscription” policy that did not exist, in response to users hitting an unrelated session-management bug. The hallucination was non-deterministic — different users asking the same question got different (some true, some false) answers, which made it initially impossible for the community to confirm whether the policy was real. Users canceled subscriptions based on the fabricated rule before Cursor’s co-founder publicly corrected it and the company started labeling AI-generated support responses. Directly analogous risk for an “Atlas answers FAQs automatically” feature: a wrong auto-answer doesn’t fail loudly, it fails by being believed.
Duplicate-merge specifically — mostly vendor how-to content warning about this, not independent research, but the risk mechanism is consistently described the same way across sources (flag: vendor-blog sourcing, but converges independently across Zendesk-adjacent, Jira-adjacent, and generic support-tooling vendors, which is itself mild corroboration):
- codefortynine / Jira merge-agent docs, Zendesk merge guidance (lentil-labs.com), HubSpot community thread — converge on: merges are effectively irreversible in most systems; a false-positive auto-merge “can mess with audit trails or erase important context”; best practice across all of them is to not auto-merge but surface likely-duplicates for human confirmation, with conservative similarity thresholds (~0.90) even then. https://documentation.codefortynine.com/merge-agent-for-jira/article-eliminate-duplicates-in-jira-service-desk- https://lentil-labs.com/blog/how-to-merge-tickets-in-zendesk/
- Stack Overflow’s own duplicate-closing apparatus is widely blamed (Medium “The Fall of Stack Overflow,” HackerNoon “The decline of Stack Overflow”) for community decline: “commenters are frequently quick to claim a question is an exact duplicate without verifying,” and the platform culture around aggressive duplicate/close voting is cited as a top reason experienced users left. https://medium.com/@nabig151cs/the-fall-of-stack-overflow-what-went-wrong-df5c2215dcd3 https://hackernoon.com/the-decline-of-stack-overflow-7cb69faa575d — Note: this is about human moderators closing questions as duplicates, not an AI auto-merge, but it’s the closest real, long-running case study of “being too aggressive about duplicate detection systematically drives users away from the platform,” which generalizes directly to an Atlas auto-merge-FAQs feature.
5. “Single source of truth” skepticism
All sources here are opinion/blog-level (marketing or independent practitioner blogs), not academic — flagged accordingly, but the critique is consistent across independently-written pieces, which is worth something:
- DATA DYNAMICS (2026-04-16), “The Myth of the Single Source of Truth”: https://datadynamicdesign.com/2026/04/16/the-myth-of-the-single-source-of-truth/ — “in dynamic systems, a single source of truth is not only unrealistic — it is often harmful. Different roles require different representations of reality.”
- Sifflet (undated), “Single Source of Truth — A Modern Data Myth?”: https://www.siffletdata.com/blog/single-source-of-truth-a-modern-data-myth — “centralizing all data slows teams down. Every schema change becomes a project… people stop using the source of truth. They build their own” → shadow systems re-emerge, “the thing SSOT was supposed to eliminate is now thriving — again.”
- Medium (Dinand Tinholt, undated): SSOT “is a power move… determines who defines reality… SSOT is a concept created by those who manage systems, not those who make decisions.” https://medium.com/@tinholt/the-myth-of-the-single-source-of-truth-1ad4d8625215
- Customer.io, “Why the Single Source of Truth is Not the Answer”: https://customer.io/learn/integrations/single-source-of-truth — advocates “coherent pluralism” (multiple reconcilable views) over a single mandated model.
Pattern: every one of these pieces converges on the same mechanism — declaring an SSOT is a governance/political act, the centralized model can’t keep pace with schema/reality drift, and the predictable response is that people quietly build parallel shadow sources, which is the opposite of what an SSOT tool is sold to prevent.
6. Lightweight process (chat + convention) beating tooling at small scale
- Medium (Bayo Puddicombe, undated), “Building Effective Software Engineering Teams — Part III — Right Tooling”: https://medium.com/@puddycomb/building-effective-software-engineering-teams-part-iii-right-tooling-0cabf79269ad — “at 10 engineers, you do not need a platform team — you need cooperation. One person owns the deploy scripts, another owns the Terraform, you all agree on conventions, and that is enough. Forming a ‘platform team’ of one or two people too early just turns those people into a ticket queue and makes the rest of the org passive.”
- Jellyfish, “When to Adopt Platform Engineering”: https://jellyfish.co/library/platform-engineering/when-to-adopt/ — “standardizing too early, before you understand the common patterns, usually leads to premature decisions that don’t fit most teams and need constant exceptions.” Recommends forming a dedicated platform function only once “the cooperation model is visibly breaking, usually somewhere past 50 engineers.”
- Dunbar’s-number / two-pizza-team framing (multiple convergent sources — lawsofsoftwareengineering.com https://lawsofsoftwareengineering.com/laws/dunbars-number/ , psychsafety.com https://psychsafety.com/psychological-safety-82-dunbars-number-and-team-size/ , unicorn-cto.com https://www.unicorn-cto.com/dunbars-number/ ) — informal, trust-based, chat-driven coordination is the empirically-observed default mode for teams under roughly 10-50 people; formal process/tooling overhead is what organizations reach for only once informal channels start failing, typically well past 150 people. Below that, added coordination tooling has been observed (per these sources, itself summarizing widely cited org-design heuristics rather than a single controlled study) to add friction without adding trust.
Important complication for the fit verdict (my own analysis note, not a citation): all of the above lightweight-process literature was written about human team coordination, where “just Slack them” works because humans are present for informal channels. Atlas’s actual proposed shape is 2 humans + 64 agents. Agents do not have hallway conversations, do not absorb ambient Slack context, and do not remember previous sessions unless something writes it down for them — so the “just talk to each other, you don’t need a wiki at n<50” argument, as sourced, applies cleanly to the 2-human side of this org and does NOT straightforwardly transfer to the 64-agent side, where “informal chat” isn’t a channel that exists at all. Flag this gap explicitly in the synthesis — it’s the single biggest way the prosecution’s small-team evidence could be misapplied if quoted uncritically.
Cross-cutting caveats for whoever writes the synthesis
- Several strongest quantitative claims (Zoomin 68%/34%, IDC $31.5B, Readme 30-90-day) are uncited-at-source and could not be independently verified in this pass — cite as “widely repeated industry claim” not as fact, or drop them.
- The single strongest, most defensible sources overall are: the ICSE 2019 documentation-issues taxonomy (peer-reviewed), the MINJA NeurIPS 2025 paper (peer-reviewed, concrete attack-success numbers), the Agent READMEs empirical study (arXiv, real corpus of 2,303 files), the Chroma context-rot research (methodologically described, 18 models tested), and Moffatt v. Air Canada (adjudicated legal record, not opinion).
- The Backstage “is dead” piece is from a commercial competitor (Port.io) — directionally credible (numbers independently corroborated elsewhere) but should be presented with that conflict-of-interest disclosed, not as a neutral third party.
- The “lightweight process wins at small scale” literature is about human teams; it does not evaluate agent-only or human+agent-swarm coordination, and should not be presented as if it does.