Alle Beiträge
Blog

How to Validate a Startup Idea Before Writing Code

Vom RawKit-Team
11 Min. Lesezeit

Dieser Beitrag ist derzeit nur auf Englisch verfügbar.

How to Validate a Startup Idea Before Writing Code

What validation actually requires

Before you write a line of code, you need answers to four questions: who has the problem, who pays to fix it, who else is trying to solve it, and what evidence backs each assumption. The cheapest way to answer all four is reading what already exists in public. A developer forum post, an app store review, a Reddit thread—these are all receipts from people who have already made decisions you need to understand. Building comes later, after the reading is done.

The four questions carry a specific shape, and the shape is the point. A founder who answers "yes, the problem is real" without a who and a what-evidence has done half a check, and the missing half is what kills the idea later when customers do not show up. The order does not matter; the completeness does.

Public record is the cheapest source. A Reddit thread where someone describes the workaround they built is one data point. An app store review calling out a missing feature is another. A job posting for a role that would not exist if the problem were already solved is another. The cost of gathering these is not measured in dollars so much as in the time it takes to read them, and that is the point. The real money gets spent later, in engineering, and it should not be spent on guesses.

A receipt for this kind of work exists in the public record: a research pass that pulled nineteen sources, ran the four questions across them, and came to $0.02. That is what a basic validation check looks like—not a weekend of work, not an outside consultant, but a few cents of compute and a focused read. The receipt sits at the lower bound, not the average. The point is that the floor is now low enough that skipping it is the more expensive choice.

Where the problem already lives

The problem exists where people are already spending money or time to solve it. Look for complaints on Reddit threads about existing tools failing. Check App Store and Google Play reviews for one-star complaints that mention specific missing features. Search Hacker News for the problem described in plain terms. Each source where people describe the same pain is one data point. The evidence lives in the language people use to describe what they need fixed.

The signal is in the verbs. When a user says "I tried X and it does not Y," that is a complaint. When they say "I built this script to do Z," that is evidence of demand. When they say "I pay for W but it has never been able to Q," that is a customer describing their existing spend and naming the gap.

A few types of source do most of the work. Community forums (Reddit, Hacker News, LinkedIn) carry complaints in plain language and let you filter by recency. App stores carry complaints from people who already paid once, which is a stronger signal than a comment from someone passing through. Funding databases (EU Funding Portal, Grants Institutions) carry the question the other way around: they show who is paying to chase a problem, which is the same problem from a different angle.

Source typeWhat it tells youWhat it does not
Reddit, Hacker News, LinkedInProblem described in the user's own wordsWhether the user would pay
App Store, Google PlayUser already paid once and is still unhappyWhether the user would switch
GitHub, arXivEngineers actively building toward the problemWhether non-engineers feel it
Amazon, EtsyPeople buying adjacent workaroundsWhether the workaround is good enough
Eurostat, OpenAlexPopulation-scale demand and prior artWhether the demand is in your segment
EU Funding Portal, Grants InstitutionsPublic money flowing into the spaceWhether private money will follow

Tools that automate this kind of pull exist. RawKit, for instance, gathers nineteen source types—Reddit, Hacker News, LinkedIn, X, App Store, Google Play, Amazon, Etsy, GitHub, Eurostat, OpenAlex, arXiv, EU Funding Portal, and Grants Institutions—and drops each one onto a research board as a card the user can open. The model never writes a quote. Every quote lands on the board copied from a page the user can open, never written by the model. That is the part that matters: a sentence in a comment is only useful if you can verify the page it came from. The number of structured card kinds in such a system is seventeen, each with named fields, so the resulting research is comparable across ideas rather than trapped in a doc.

Without a tool, the practice still works. Open Reddit, open the App Store, open GitHub, copy the relevant posts into a shared doc. The lift is higher and the structure is worse, but the underlying habit—reading what people wrote before you build what you think they want—is the same.

Following the money already being spent

Someone pays when they hand over money or dedicate significant time. Look for where people are already spending: subscription revenue figures on competitor websites, open-source star counts on GitHub, funding announcements in your space. These numbers are public. They tell you whether a market exists, how large it is, and whether people are spending actual money or just time. A market with spending is a market worth entering. One without is a hypothesis.

Time is a form of payment. A user who has built their own workaround has already paid in hours, which is the same kind of demand as someone paying in dollars, just at a different rate. Founders who only count revenue miss the unpaid side, and the unpaid side is often the larger one. A problem that costs people a Sunday afternoon every month is a real market, even if no one has built the paid version yet.

The strongest signal is actual spend by adjacent products. If competitors exist and each publishes pricing on a website, that is a market. If a competitor's GitHub repo is heavily starred, that is a market. If a funding announcement landed recently, that is a market with capital behind it. None of these prove the founder's specific idea will work. They prove the problem is being paid for in some form, somewhere, by someone other than the founder.

Confidence is computed, not claimed. Each piece of evidence feeds a number from 0 to 100, and the number is arithmetic, not a model's opinion. A claim backed by one Reddit thread scores lower than a claim backed by several threads plus a competitor's pricing page plus a funding announcement. The score moves when the evidence moves, and the score is comparable across the four questions in a way that feelings are not. A research tool that only returns good news has told you nothing; the score is what forces the question "where is the evidence weakest" to surface before the build starts.

Evidence kindWhat it does to the confidence scoreWhy
Multiple users describe the same problem recentlyMoves score up sharplyRecency plus repetition equals real demand
Competitor publishes pricing and users complain about itMoves score up moderatelyExisting spend is confirmed, gap is named
One user describes the problem long agoMoves score up slightlySpecificity without recency is weak
Trend report citing aggregate "interest"Moves score sidewaysAttitudes do not predict behavior
User describes the problem but has not paid for any workaroundMoves score up less than a paying userTime-only payment is a weaker signal than money

Recent, specific, and tied to a name

A forum post from last month with a specific complaint beats a trend report from years ago. A recent Reddit thread where someone paid for a workaround beats an old survey. The recency and specificity of evidence matter more than volume. One person describing exactly the problem your idea solves, made this week, is worth more than a hundred vague expressions of interest. The evidence should point to a specific behavior, not a general attitude.

A complaint without a product name attached to it is closer to a feeling than a signal. "This is so annoying" with no referent is not evidence. "I pay for Calendly and it still does not handle round-robin for sales calls" is evidence, because it names a paid product, a price, and a specific missing feature. The specificity is the proof that the user has thought about this, and the recent date proves the feeling has not cooled.

The recency window depends on the market. A consumer app that turns over quickly needs evidence from the last quarter. A B2B workflow tool in a regulated industry might still find older complaints useful if the regulations have not changed. The point is not a single rule but a habit: the recency field on every piece of evidence is read before it is trusted. A tool that pulls nineteen sources and timestamps each one makes this easy. A shared spreadsheet with a "date" column does the same job, slower.

Specificity has layers. The first is the named product or workaround: "I tried X and it broke." The second is the named outcome: "and now I have to do Y by hand every week." Without the second, the evidence is sympathy for the user, not a description of a job to be done. Sympathy is fine. A job to be done is what the build needs.

Commit, drop, or keep reading

You commit when you can answer all four questions with evidence, not opinions. You drop when two or more of the four questions have no public record supporting them. Between those points, the answer is a third state: more reading. The goal is not certainty—it is a confidence score above a threshold you set for yourself. The number is yours to choose, but it must come from evidence, not feeling.

The threshold is a number the founder sets before the reading starts, not after. A pair of reasons drives this. The first is that the temptation to move the goalposts at the end is strong; every founder wants their idea to be the one that survives, and that want distorts the read. The second is that the threshold becomes a contract the founder can show to a co-founder, an investor, a partner, or themselves later when the evidence has not changed but the feeling has.

A common threshold is a high score across all four questions, with no individual question sitting low. That is not a universal rule. Some founders set higher bars for problems they have lived. Some set lower bars for problems they have personally paid to solve. The point is that the number is set before the evidence comes in, and the decision to commit or drop is made against that number, not against a feeling at the end of the read.

A drop is not a failure. A drop in the first weeks saves the months a build would have cost. A drop also means the founder now has a tested hypothesis about which adjacent idea is closer to a market, which is information they did not have before the read. The third state—more reading—is the most common outcome of an honest first pass, and it is the state that produces the most founders who actually ship, because by the time they commit, the build is a formality rather than a leap.

The minimum viable research pass

A real validation pass requires reading across at least 19 sources. Reddit, Hacker News, LinkedIn, app stores, and funding databases each answer different parts of the four questions. Email-only signup with $5 in free credits is enough to start—many founders run their first check before committing anything. The cost can be as low as $0.02. No credit card on file. No build required. A spreadsheet and a saved Reddit thread work too.

Nineteen sources is a count, not a prescription. The principle is broader: the read has to span at least one source from each of the four buckets—problem, payment, competition, evidence—so that the founder is not making a decision on a single angle. A research pass that only looks at Reddit will over-index on developer complaints. A pass that only reads app store reviews will miss the B2B angle. The breadth is the point.

The signup itself is the lightest barrier that still works. The first 50 signups get $5 in credits; an email and a verification link is all that is required. No credit card. The point of that design is to keep the cost of an honest first attempt at near zero, so that the only thing standing between a founder and a tested idea is the act of reading. If the founder never signs up, the practice still works—the spreadsheet, the saved Reddit thread, the manual screenshot of a competitor's pricing page all do the same job at higher cost in time.

The receipt is real. $0.02 covers the cost of a research pass that pulled the nineteen sources, ran the four questions across them, and produced a confidence score per claim. That is not a marketing number. It is what one actual run cost, and the cost is what the floor looks like when the loop is automated. A founder doing the work by hand will spend more in time and less in dollars, but the structure of the work is the same: read, score, decide.

The cost of skipping the pass is harder to measure but easier to feel. Three months of engineering on a problem nobody pays to solve produces a working product and a quiet launch. The $0.02 read would have caught it in the first week. The math is not subtle.

Mehr aus dem Blog