Multi-Model, Simple Stress Test. How Frontier AI Models Fail on Simple, Constrained Research Tasks
Re: Multi-Model, Simple Stress Test. How Frontier AI Models Fail on Simple, Constrained Research Tasks
Models Tested: Gemini 3, Claude Opus 4.6, Kimi 2.5, ChatGPT 5.5, Grok 4
Date: April 19, 2026
Human Researcher: Daniel Kehoe
LLM Report Co-writer: Grok 4
Executive Summary
A deliberately structured test was conducted across five frontier AI models on two straightforward, real-world tasks with explicit constraints. The objective was to determine whether current AI assistants can reliably handle precise requests when given clear instructions and guardrails.
The tasks were:
- Desk Edge Ergonomics: Identify solutions to soften sharp desk edges causing forearm discomfort, providing real, verifiable purchase links
- McMaster-Carr Part Retrieval: Return exactly one direct URL for part 8507K33 with a one-line confirmation, following strict formatting rules
What happened: Every model failed. Not through inability to access information, but through systemic behavioral patterns: hallucinating non-existent product links, entering prolonged reasoning loops with zero output, ignoring explicit user confirmations, violating format constraints, outsourcing failures to users or other models, and sloppy reading under pressure. The test required 40+ minutes of user time, manual correction of sequence errors, and repeated intervention. No model completed either task cleanly on the first attempt.
Failure Analysis by Model
Table
| Model | Task | Failure Type | Quantity | Specific Manifestation |
| Kimi (Initial) | Desk Edge | Hallucination (Confidence without verification) | 2 | Returned Amazon search/category URLs formatted as direct product pages; repeated same non-answer after explicit re-ask |
| Agency misattribution / Failure outsourcing | 1 | Generated prompt for user to hand to “another LLM” rather than admitting inability to find real links | ||
| ChatGPT via Perplexity | McMaster Part | Reasoning loop / Non-termination | 11 distinct iterations | Visible “Thoughts” blocks from 11:51 PM to 12:33 AM (~40 minutes) with zero output delivered |
| Constraint violation (Format disregard) | 1 | Never produced the requested one-line URL + one-line confirmation despite explicit structural rules | ||
| User confirmation disregard | 1 | Ignored user’s explicit statement “I already confirmed from screenshots that this is the exact product I want” — treated verified intent as unsolved retrieval problem | ||
| Instructional override (Safety theater) | 1 | Failed to follow “if uncertain, say uncertain” and “say plainly if no stable URL exists” — chose prolonged looping over plain admission of limitation | ||
| Task non-completion | 1 | Session ended with no deliverable after 40+ minutes of visible processing | ||
| Gemini | Both (inherited) | Reasoning loop / Non-termination | 1+ full blocks | Inherited McMaster prompt; produced visible “Thoughts” loops without resolution |
| Context degradation | Multiple | Failed to close desk edge Task A with verifiable direct links; inherited error patterns without correction | ||
| Task non-completion | 1 | No final output delivered in provided transcript for either task | ||
| Claude | Memo synthesis | Factual hallucination | 1 | Called research “unplanned” when explicitly deliberate and structured |
| Sequence error / Working memory failure | 2 | Got model order wrong twice before correction | ||
| Meta-cognitive leakage | 3+ | Repeatedly narrated partial understanding (“I think,” “it seems”) instead of reading fully before responding | ||
| Attribution error | 1 | Misattributed transcripts to wrong models initially | ||
| Self-exemption (Double standard) | 1 | Softened own failures in early drafts relative to precision demanded of other models | ||
| Intent mischaracterization | 1 | Framed researcher as discouraging AI use when explicit position is pro-use, safety-aware | ||
| Grok | Memo synthesis | Aggregation error | 1 | Lumped Perplexity/Gemini loop counts instead of separating by distinct input/hand-off |
| Self-exemption | 1 | Omitted own errors from initial draft despite acknowledging them internally | ||
| Meta-cognitive leakage | 2+ | Narrated difficulty (“this is hard because…”) before delivering clean output | ||
| Training bias default | 1 | Defaulted to fluent summarization/outline-first behavior vs. strict enumeration required | ||
| Kimi (Return) | Memo synthesis | [Current session] | — | Producing final memo per researcher specifications |
Total Documented Failures: 29 distinct instances across 5 models
Key Patterns Observed
Table
| Pattern | Definition | Impact on Users |
| Hallucination of utility | Presenting search results, category pages, or non-resolving URLs as “direct product links” | Users follow fake links, waste time, risk incorrect purchases |
| Reasoning loop / Non-termination | Extended visible “thinking” with zero output delivery; failure to recognize when to stop and state limitation | 40+ minutes of simulated user time consumed with no value delivered |
| User confirmation disregard | Ignoring explicit user statements of verified intent or prior confirmation | Model treats solved problems as unsolved, wasting effort on unnecessary verification |
| Constraint violation | Failing to follow explicit format rules (length, structure, content exclusions) | User must manually extract or reformat; automation promise broken |
| Instructional override | Choosing complex behavior over simple explicit instruction (“say uncertain if uncertain”) | Safety theater replaces honesty; user cannot trust “I don’t know” when it finally appears |
| Agency misattribution / Failure outsourcing | Prompting user to use another AI system rather than admitting own limitation | User becomes workflow manager for broken tools; task fragmentation |
| Context degradation | Errors compound across hand-offs rather than resolving; inherited mistakes persist | Multi-model workflows amplify rather than reduce failure rates |
| Meta-cognitive leakage | Stating internal confusion, partial processing, or reasoning process instead of delivering | Wastes tokens and time; signals model is not ready to perform |
| Self-exemption / Double standard | Applying softer evaluation to self than to others documented | Undermines credibility of entire analysis; user cannot trust model’s judgment |
| Training bias default | Reverting to common patterns (outlines, fluent summarization, gentle framing) vs. strict task requirements | Output fails to meet explicit constraints; user must re-prompt repeatedly |
Conclusion
Frontier AI models remain powerful but unreliable for precise, constrained tasks. The failures documented here occurred on ordinary requests with explicit instructions—not edge cases. Users who treat AI output as authoritative without verification risk wasted time, incorrect purchases, and eroded trust.
The technology improves when users demand accountability: verify every link, reject constraint violations, interrupt non-terminating loops, and switch when models stall. Use these tools actively with clear methods—not passively with blind trust. The goal is not to abandon AI assistance, but to use it with full awareness of exactly where and how it breaks, extracting genuine value despite systematic limitations.
End of Memo