A concise reading guide to the original source, the discussion's attention signal, and the claims that still need verification.
The short version
On-device models that know when they're wrong: every answer carries a confidence score for cloud handoff. - cactus-compute/cactus-hybrid
Why it drew attention
The collected source record shows 154 points, 35 comments, and published 2026-07-22.
Points and comments measure interest at one moment; they do not verify the linked article or settle the debate.
What to question
Identify the author's central claim, then look for primary documents, reproducible evidence, corrections, and important missing context.
Distinguish statements in the original article from interpretations added in the comment thread.
Read it in this order
Open the linked source first and note its evidence. Then read the Hacker News comments for counterexamples, corrections, and additional references. Recheck any important conclusion against a primary source.
Bottom line
This discussion is useful as a map of questions and reactions, not as independent proof. The most reliable takeaway comes from comparing Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong with the primary evidence it cites.
Read the original discussion
Open the original source, look for primary evidence, and consider corrections or later developments. Community voting provides context but does not verify a story's claims.
At a glance
Hacker News points are a time-specific attention signal, not a software adoption metric.
Sources and methodology
This guide uses public records collected on 2026-07-23. Source completeness is High; this label is not a product rating.
View the supporting source data
- Comments at collection: 35 — official record
- Points at collection: 154 — official record
- Author: HenryNdubuaku — official record
- Discussion text: Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score between 0 and 1. Developers can accept the on-device when it's high, hand off to a bigger cloud model when it's low. By routing only 15-35% of queries to Gemini 3.1 Flash-Lite, Gemma-4-E2B matches Gemini 3.1 Flash-Lite on most benchmarks. - ChartQA: 15-20% - LibriSpeech: 25-30% - MMBench, GigaSpeech, MMAU: 30-35% - MMLU-Pro: 45-55% We were always frustrated by the routing signals hybrid apps rely on: asking the model to — official record
- Hacker News item: 49.0M — official record
- Published: 2026-07-22 — official record
- Story title: Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong — official record
- Original source summary: On-device models that know when they're wrong: every answer carries a confidence score for cloud handoff. - cactus-compute/cactus-hybrid — official record
Evidence limitations
- Hacker News points and comments are time-specific attention signals, not independent verification of the linked claims.
- Official metadata was collected as source evidence; no independent installation, benchmark, security audit, or practical product evaluation is claimed.
Data note: This automated workflow updated the guide on 2026-07-23 from public Hacker News story records. Metrics can change after collection. No independent product testing or endorsement is claimed.