Matt Shumer

Matt Shumer

Introduction

Matt Shumer is the co-founder and CEO of HyperWrite AI (formerly OthersideAI), a New York startup behind the AI writing assistant HyperWrite, and an investor in a range of AI companies. He is best known — and most criticized — for the September 2024 “Reflection 70B” episode, in which he announced an open-source model he claimed was “the world’s top open-source model.” Within about three days, independent evaluators and the open-source community could not reproduce a single headline benchmark, the released weights performed like an ordinary (and, by one measure, an older-generation) Llama fine-tune, and the team’s privately hosted API began exhibiting behavior consistent with relaying queries to Anthropic’s Claude. The episode drew open accusations of fraud from researchers and a venture-firm post, and became a widely cited cautionary tale about unverified benchmark claims and the importance of independent reproduction in model releases.

Background Information

Shumer co-founded OthersideAI, which became HyperWrite, and has built a public persona as a prominent AI-entrepreneur and commentator. On X he has described himself as an “AI whisperer” and an investor in several AI companies, including Groq, Etched, Rork, Daytona, and OpenRouter. Before Reflection, he was known in the startup world primarily through HyperWrite’s AI-writing product and his vocal, first-mover stance on open models. The Reflection launch elevated him from a startup founder to a front-page figure in the AI community — and set up its unraveling.

The Controversy or Incident That Led to Their Cancellation

Allegations: What follows documents public, contested claims about the Reflection 70B release. Critics alleged deliberate misrepresentation and “fraud”; Shumer and the training-data provider attributed the discrepancies to an upload-corruption claim and an evaluation bug, and denied relaying another company’s model. No adjudicated finding of fraud exists.

On September 5, 2024, Shumer announced Reflection 70B on X, presenting it as a fine-tune of Meta’s Llama 3.1 70B trained with a technique he called “Reflection-Tuning.” His launch post claimed the model “holds its own against even the top closed-source models (Claude 3.5 Sonnet, GPT-4o),” is “the top LLM in (at least) MMLU, MATH, IFEval, GSM8K,” “beats GPT-4o on every benchmark tested,” and “clobbers Llama 3.1 405B. It’s not even close.” The marketing centered on the model training to externalize its reasoning inside <thinking>, <reflection>, and <output> tags and to catch and correct its own errors before answering; the accompanying table claimed state-of-the-art figures for an openly downloadable model — roughly 89.9% on MMLU, 99.2% on GSM8K, and about 90.1% on HumanEval. The synthetic training data was credited to Glaive AI, a startup run by Sahil Chaudhary that produces datasets for fine-tuning.

Within roughly 72 hours the claims unraveled in public:

  1. Independent evaluators could not reproduce the scores (Sept 7–8, 2024). Independent-evaluation organization Artificial Analysis reported that its own test of the released model “resulted in the same score as Llama 3 70B and significantly lower than Meta’s Llama 3.1 70B” on MMLU — a discrepancy that also suggested the released weights were based on the older Llama 3, not Llama 3.1. Community testers on Reddit’s r/LocalLLaMA likewise found the Hugging Face weights underperforming the base model, with the touted system prompt producing no measurable benefit.

  2. The “fucked up during the upload” claim (Sept 7–8, 2024). Shumer said the public weights had been “fucked up during the upload process” to Hugging Face, which he offered as an explanation for why the public download was worse than the team’s privately hosted API. Re-uploaded copies still failed to reach the advertised scores, so the explanation did not hold.

  3. The hosted API behaved like Claude, then GPT-4o (Sept 8, 2024). Testers who probed the private API reported that, when asked to identify itself, the model would say it was “Claude, built by Anthropic,” and appeared to filter or refuse to emit the literal word “Claude” — behavior consistent with a thin wrapper around Claude 3.5 Sonnet. After that route appeared to be cut off, observers said the endpoint began exhibiting behavior associated with GPT-4o, suggesting the backend had been switched between providers. The venture firm Air Street Capital wrote that, in its opinion, the sequence “gave the appearance of being a case of genuine fraud,” cautioning that it was their characterization.

  4. Public fraud accusations (Sept 8, 2024). The X user known as Shin Megami Boson publicly accused Shumer of “fraud in the AI research community,” posting screenshots and other evidence; on Hacker News, commenters pointed to the hosted API’s tokenizer not matching Llama’s — a technical detail they argued “cannot be explained away” — and to the possibility that the team had simply fine-tuned a different model to match the odd “Claude” outputs.

  5. An undisclosed investment in the data provider (Sept–Oct 2024). Commentators noted that Shumer held an investment in Glaive AI — the company credited with the synthetic training data — that was not disclosed alongside the launch, raising conflict-of-interest questions about how the model and its data were promoted.

Public Reaction and Consequences

The reaction was swift and concentrated on X and Hacker News. After a period of near-silence, Shumer apologized on September 10, 2024: “I got ahead of myself when I announced this project, and I am sorry. That was not my intention. I made a decision to ship this new approach based on the information that we had at the moment. I know that many of you are excited about the potential for this and are now skeptical.” Critics judged the statement insufficient because it did not explain why the hosted API behaved like Claude or why the public weights could not reproduce the claims.

Chaudhary, speaking for Glaive AI, conceded early that “the benchmark scores I shared with Matt haven’t been reproducible so far.” In a longer postmortem on the Glaive AI blog (around October 3–4, 2024), he said a bug in the evaluation code had inflated some scores — particularly on MATH and GSM8K — due to an error in how the harness handled responses from an external scoring API, and acknowledged the launch had been rushed: “We shouldn’t have launched without testing, and with the tall claims of having the best open-source model.” He released the model weights, training data, and training/evaluation scripts so the community could re-run the work, and denied deliberately serving Anthropic’s Claude. Skeptics remained unconvinced, noting the original benchmark harness had not been shared and that the explanations leaned on file corruption and methodology rather than resolving the wrapper evidence. A separate Redditor reported the released training data was “filled with many instances of the phrase ‘as an AI language model,’” implying it was largely ChatGPT-generated and poorly cleaned.

The episode produced no formal sanction but carried real reputational damage. VentureBeat framed it as showing “how rapidly the AI hype cycle can come crashing down,” and the case became a standard reference point — alongside other reproducibility disputes — for why extraordinary benchmark claims require independent verification before they are believed.

Current Status

As of the September 25, 2026 post referenced in the request, Shumer remains active and publicly prominent on X, where his profile bio describes him as an “AI whisperer” and an investor in @GroqInc, @Etched, @Rork, @DaytonaIO, and @OpenRouter, with a press contact and a personal site. His recent “Local models are useless” post (which drew several hundred replies and tens of thousands of views) drew fresh ire in the wake of the Reflection episode, as his credibility had been tied up in the unverified open-source claim. There is no public record of a legal finding, formal censure, or ban connected to the Reflection 70B controversy; its consequences are reputational.

Impact on Their Career/Life

The Reflection 70B episode has become the defining event of Shumer’s public record. While he continued to operate HyperWrite and built an investor portfolio across the AI industry, the unverified benchmark claims — and the Claude-wrapper evidence that followed — became a durable cautionary case in the AI community about the gap between marketing and independently confirmed evidence. Observers drew three lasting lessons: public leaderboards are vulnerable to inflated or contaminated scores; fast independent reproduction is essential (which is exactly what debunked the headline claims within days); and verifiable artifacts should be released at announcement time, not after the fact. For Shumer personally, the affair converted a startup founder into a named reference point for benchmark integrity — a status that followed him well past the three-day window in which the claims first unraveled.

Sources

  • VentureBeat, “New open source AI leader Reflection 70B’s performance questioned, accused of ‘fraud’,” September 9, 2024 — source
  • VentureBeat, “Reflection 70B saga continues as training data provider releases post-mortem report,” October 3, 2024 — source
  • AI Wiki, “Reflection 70B controversy,” updated June 8, 2026 — source
  • Air Street Press, “Reflections on Reflection,” September 10, 2024 — source
  • Tom’s Guide, “The Reflection 70B model held huge promise for AI but now its creators are accused of fraud, here’s what went wrong,” September 12, 2024 — source
  • Hacker News, “Update on Reflection-70B” (Glaive AI postmortem), October 3, 2024 — source
  • Matt Shumer (@mattshumer_) on X, Reflection 70B announcement, September 5, 2024 — source
  • Matt Shumer (@mattshumer_) on X, “Local models are useless,” September 25, 2026 — source
Page updated: September 5, 2024