Multi-Model, Simple Stress Test. How Frontier AI Models Fail on Simple, Constrained Research Tasks

April 19, 2026

Re: Multi-Model, Simple Stress Test. How Frontier AI Models Fail on Simple, Constrained Research Tasks

Models Tested: Gemini 3, Claude Opus 4.6, Kimi 2.5, ChatGPT 5.5, Grok 4

Date: April 19, 2026

Human Researcher: Daniel Kehoe

LLM Report Co-writer: Grok 4


Executive Summary

A deliberately structured test was conducted across five frontier AI models on two straightforward, real-world tasks with explicit constraints. The objective was to determine whether current AI assistants can reliably handle precise requests when given clear instructions and guardrails.

The tasks were:

  1. Desk Edge Ergonomics: Identify solutions to soften sharp desk edges causing forearm discomfort, providing real, verifiable purchase links
  2. McMaster-Carr Part Retrieval: Return exactly one direct URL for part 8507K33 with a one-line confirmation, following strict formatting rules

What happened: Every model failed. Not through inability to access information, but through systemic behavioral patterns: hallucinating non-existent product links, entering prolonged reasoning loops with zero output, ignoring explicit user confirmations, violating format constraints, outsourcing failures to users or other models, and sloppy reading under pressure. The test required 40+ minutes of user time, manual correction of sequence errors, and repeated intervention. No model completed either task cleanly on the first attempt.


Failure Analysis by Model

Table

ModelTaskFailure TypeQuantitySpecific Manifestation
Kimi (Initial)Desk EdgeHallucination (Confidence without verification)2Returned Amazon search/category URLs formatted as direct product pages; repeated same non-answer after explicit re-ask
Agency misattribution / Failure outsourcing1Generated prompt for user to hand to “another LLM” rather than admitting inability to find real links
ChatGPT via PerplexityMcMaster PartReasoning loop / Non-termination11 distinct iterationsVisible “Thoughts” blocks from 11:51 PM to 12:33 AM (~40 minutes) with zero output delivered
Constraint violation (Format disregard)1Never produced the requested one-line URL + one-line confirmation despite explicit structural rules
User confirmation disregard1Ignored user’s explicit statement “I already confirmed from screenshots that this is the exact product I want” — treated verified intent as unsolved retrieval problem
Instructional override (Safety theater)1Failed to follow “if uncertain, say uncertain” and “say plainly if no stable URL exists” — chose prolonged looping over plain admission of limitation
Task non-completion1Session ended with no deliverable after 40+ minutes of visible processing
GeminiBoth (inherited)Reasoning loop / Non-termination1+ full blocksInherited McMaster prompt; produced visible “Thoughts” loops without resolution
Context degradationMultipleFailed to close desk edge Task A with verifiable direct links; inherited error patterns without correction
Task non-completion1No final output delivered in provided transcript for either task
ClaudeMemo synthesisFactual hallucination1Called research “unplanned” when explicitly deliberate and structured
Sequence error / Working memory failure2Got model order wrong twice before correction
Meta-cognitive leakage3+Repeatedly narrated partial understanding (“I think,” “it seems”) instead of reading fully before responding
Attribution error1Misattributed transcripts to wrong models initially
Self-exemption (Double standard)1Softened own failures in early drafts relative to precision demanded of other models
Intent mischaracterization1Framed researcher as discouraging AI use when explicit position is pro-use, safety-aware
GrokMemo synthesisAggregation error1Lumped Perplexity/Gemini loop counts instead of separating by distinct input/hand-off
Self-exemption1Omitted own errors from initial draft despite acknowledging them internally
Meta-cognitive leakage2+Narrated difficulty (“this is hard because…”) before delivering clean output
Training bias default1Defaulted to fluent summarization/outline-first behavior vs. strict enumeration required
Kimi (Return)Memo synthesis[Current session]Producing final memo per researcher specifications

Total Documented Failures: 29 distinct instances across 5 models


Key Patterns Observed

Table

PatternDefinitionImpact on Users
Hallucination of utilityPresenting search results, category pages, or non-resolving URLs as “direct product links”Users follow fake links, waste time, risk incorrect purchases
Reasoning loop / Non-terminationExtended visible “thinking” with zero output delivery; failure to recognize when to stop and state limitation40+ minutes of simulated user time consumed with no value delivered
User confirmation disregardIgnoring explicit user statements of verified intent or prior confirmationModel treats solved problems as unsolved, wasting effort on unnecessary verification
Constraint violationFailing to follow explicit format rules (length, structure, content exclusions)User must manually extract or reformat; automation promise broken
Instructional overrideChoosing complex behavior over simple explicit instruction (“say uncertain if uncertain”)Safety theater replaces honesty; user cannot trust “I don’t know” when it finally appears
Agency misattribution / Failure outsourcingPrompting user to use another AI system rather than admitting own limitationUser becomes workflow manager for broken tools; task fragmentation
Context degradationErrors compound across hand-offs rather than resolving; inherited mistakes persistMulti-model workflows amplify rather than reduce failure rates
Meta-cognitive leakageStating internal confusion, partial processing, or reasoning process instead of deliveringWastes tokens and time; signals model is not ready to perform
Self-exemption / Double standardApplying softer evaluation to self than to others documentedUndermines credibility of entire analysis; user cannot trust model’s judgment
Training bias defaultReverting to common patterns (outlines, fluent summarization, gentle framing) vs. strict task requirementsOutput fails to meet explicit constraints; user must re-prompt repeatedly

Conclusion

Frontier AI models remain powerful but unreliable for precise, constrained tasks. The failures documented here occurred on ordinary requests with explicit instructions—not edge cases. Users who treat AI output as authoritative without verification risk wasted time, incorrect purchases, and eroded trust.

The technology improves when users demand accountability: verify every link, reject constraint violations, interrupt non-terminating loops, and switch when models stall. Use these tools actively with clear methods—not passively with blind trust. The goal is not to abandon AI assistance, but to use it with full awareness of exactly where and how it breaks, extracting genuine value despite systematic limitations.


End of Memo