What the verdicts and percentages mean
A verdict like “PRETTY STUPID. 71% stupid” looks like a gut reaction, but it's built in two steps. First an AI model built for fast, structured judgments answers a set of narrow questions about your text. Then fixed rules in our code turn those answers into the verdict. The AI never writes the verdict directly, and the same answers always produce the same result. This page explains the second step in full, because a number is only useful if you know what it's made of.
Is My Idea Stupid?: the ingredients
For every idea, the model rates how strongly each of these is true. The percentages are how much each one counts in the overall result:
- Need (30%): does it solve a problem people actually have? The biggest single factor, on purpose.
- Originality (20%): is it meaningfully different from what already exists?
- A plausible customer (15%): is there a specific kind of person or business who would use it?
- Money (15%): is there a realistic way it makes money?
- Easy to understand (10%): would a normal person get it quickly?
- Ease of execution (10%): how hard would it be to build and run?
Two more questions act as brakes rather than ingredients:
- Does it sound absurd? This only hurts when the fundamentals (need, money and customer) are also weak. A weird-sounding idea with real demand isn't punished just for being unusual. A weird-sounding idea with no demand is punished hard.
- Could it be illegal or harmful? This pulls the result down sharply, whatever else is true.
If the text doesn't read like an idea at all, for example a question or a random sentence, you get a polite request to describe something you want to make instead of a verdict.
The five verdicts
The ingredients combine into one “stupidity” value between 0 and 100%, which picks the verdict:
- GLORIOUSLY STUPID: 85% stupid or more.
- PRETTY STUPID: about 65 to 85% stupid.
- QUESTIONABLE: about 45 to 65% stupid.
- PROBABLY NOT STUPID: about 55 to 72% not stupid.
- SURPRISINGLY NOT STUPID: more than about 72% not stupid.
Notice the switch: for the three bad-leaning verdicts the number shows how stupid, and for the two good ones it shows how not stupid. The number always describes the verdict you got, so it reads naturally. It isn't a probability of success, and it isn't a grade. It says how firmly the result leans: 88% stupid is a confident thumbs-down, 57% not stupid is closer to a shrug.
The scores out of ten
Under the verdict you'll see five scores: Need, Easy to understand, Money potential and Originality, where higher is better, and Execution difficulty, where higher is harder. They're the same ratings that went into the verdict, rescaled to 0–10. The customer question counts in the verdict but has no score of its own.
The scores are often more useful than the verdict. Two ideas can both be “Questionable” for completely different reasons: one is clear and easy but nobody needs it, the other is badly needed and nearly impossible to build.
“Biggest problem” and the roast
The one-line “biggest problem” names the weakest specific area: weak need, unclear customer, weak path to money, already exists, hard to execute, or unclear value. Possible harm always comes first. On good verdicts the same line says “biggest risk” instead, because even a sound idea has a weakest point. If nothing is notably weak you get “No obvious dealbreaker. Now do the hard part.”
The roast is picked from a set of pre-written lines matched to the verdict or the biggest problem. It isn't generated on the spot, and the same idea always gets the same roast, so refreshing won't get you a nicer one.
“Honestly? It's torn on this one” appears when the model's answers were unsure overall. Treat those verdicts as a coin with a slight bias.
Social Verdict scores
Our sister site Social Verdict uses the same approach for messages and posts, with its own questions. Two of its checks:
- Should I send this? scores how likely you are to regret sending a message, from 0 to 100. The signals are hostility, passive-aggressive jabs, regret risk (written in anger, late at night…), neediness, oversharing, how easy it is to misread, shouty formatting and whether it makes a clear point. 75 or more is DO NOT SEND IT, 55–74 is SLEEP ON IT, 30–54 is PROBABLY FINE and under 30 is SEND IT. On the two safe verdicts the percentage is shown the other way round, as “safe to send”.
- Is this cringe? scores secondhand embarrassment from 0 to 100, using try-hard energy, humblebragging, emoji and hashtag load, forced slang, “thought leader” voice, oversharing, fishing for likes and a lack of self-awareness. 75 or more is MAXIMUM CRINGE, 50–74 PRETTY CRINGE, 25–49 MILD CRINGE and under 25 NOT CRINGE.
These scores are half the weighted average of all the signals and half the average of the two strongest ones. That way one serious problem, like a single brutal insult in an otherwise calm message, can't be averaged away by everything that's fine. If a message sounds like it involves a real safety issue, these checks don't score it at all and show a short support note instead.
What the numbers can't tell you
Every verdict is a snap judgment of the words you typed. The judge doesn't research your market, know your customers, check what launched last month, or know anything about you. It works best in English and on one idea at a time. Plenty of ideas that sounded silly became real companies, and plenty of sensible ones went nowhere. Use the verdict as a fast first reaction and a list of questions. For the slower version, read how to pressure-test a business idea.