Documentation
Version 1.1 · Last updated 2026-09-06

RawKit Documentation

Version 1.1 · Last updated 6 September 2026 · RawKit is in open beta

RawKit is a research workspace for founders. You describe an idea; research agents read what your market already said — communities, app-store reviews, marketplaces, competitor pricing pages, official statistics, public funding databases — and put what they found on an infinite canvas as structured cards, with every quote copied verbatim out of a page the run actually fetched.

These docs explain the parts you need to drive it: what the objects on the board are, how confidence is calculated, which agent to reach for, what each source costs, and what every switch in Settings does.


Start here

Signup is an email address and a verification link. There is nothing to install and no API key to obtain.

  1. Create a board. A new board is seeded with the Research stage frame and an empty Idea card inside it.
  2. Fill in the Idea card. One-liner, problem, solution, who it is for. Every later request reads it, so this is the cheapest thing you can do well.
  3. Run research. Open the chat and type /saas, /app or /physical_product followed by what you want. The agent works on the board while you watch — cards appear, evidence attaches, connectors get drawn.
  4. Read the gaps, not just the findings. Anything that came back with no evidence is your list of what to check next. A research tool that only returns good news has told you nothing.

The Playbook panel always names one next action and what would finish it. It suggests; it never blocks.

Core concepts

ConceptWhat it means in RawKit
WorkspaceYour account's container. Settings, balance, connected accounts and source switches all live at this level and apply to every board in it.
BoardOne infinite canvas: one idea, one project. Cards, connectors, files and conversations live here.
Object (card)Anything on the board — a note, a table, a file, a chat, an evidence quote, a structured artifact.
Structured artifactA card with named fields instead of free prose: a pain map, a competitor table, a hypothesis, a funding call. Comparable across boards because the fields are the same every time.
Evidence cardA verbatim quote with its url, author, date, stance and strength. The unit of proof.
ConnectorA line between two cards. Labelled supports or refutes, it feeds the confidence calculation.
ConfidenceA number from 0–100 computed from the evidence attached to a claim. Arithmetic, never a model's opinion.
GateA condition the playbook checks before it calls a step done — three independent sources, evidence on both sides, a filled field. Shown, never enforced.
ContextThe cards you select before sending a request. That selection is what the agent reads.
RunOne execution of one agent. Durable, journaled, budgeted, and inspectable while it works.
Agent modeWhich agent handles the request, and therefore which tools and which prompt it has.

Two rules run through all of it, and they are the reason the rest of the design looks the way it does:

  • The model never writes a quote. It names a page it has already read; RawKit copies the words out of that page along with the url and the author. A quote that does not appear in anything the run fetched cannot reach the board.
  • Confidence is computed, not claimed. The score comes from the evidence links you can see. A claim with nothing behind it reads zero and says so.

The board

The canvas is a real canvas first: everything below works with no agent involved.

ObjectWhat it is
NoteSticky note with rich text — headings, lists, checkboxes, code, links.
TextA free label on the canvas.
ShapeFour basic shapes, for grouping and annotation by hand.
FrameA container. Stage frames are how the board keeps the shape of your process.
TableEditable grid. Any table can be turned into a graph in place.
ChecklistTodo items the agent can also tick.
GraphA chart, usually built from a table on the board.
FileAn uploaded PDF, spreadsheet or document, with a viewer and an extracted-text drawer.
ImageAn upload or a generated image, zoomable.
VideoAn upload or a generated clip, with a player.
ChatA conversation pinned to the board as a card you can point at and link.
EvidenceA verbatim quote card with provenance and a supports/refutes badge.
MetricA single number with a unit and a delta.
Structured artifactOne of the 17 kinds in the next section.

Frames are for stages. Each of the five playbook stages gets one frame, and you create them — agents are refused when they ask for one. That keeps the board's left-to-right shape yours.

You do not place cards, and neither does the model. Coordinates were removed from the agent-facing tools: the model says what a thing is and what it relates to, and a placement engine decides where it goes. An artifact's home is the frame already holding most of the things it references, so the layout ends up explaining the reasoning. Tidy board re-plans the whole layout in one transaction and reports what it did — and it only ever runs when you ask, because rearranging a board automatically would move work under someone's cursor.

Structured artifacts

Seventeen card kinds with named fields. This is what makes two ideas comparable: a pain map from March and one from today line up field for field.

StageKinds
Researchidea, pain-map, competition, persona, market
Validatehypothesis, experiment
Strategypositioning, pricing, gtm, funding, ci
Marketingsocial-strategy, post
Analysisinsight
Planningokr, sprint

An agent fills fields one at a time, and an empty field stays visibly empty rather than being filled with something plausible. That hole is the product working correctly.

Evidence and computed confidence

Every claim-shaped card — an idea, a hypothesis, an insight, a pain — carries a confidence badge. The number is arithmetic over the evidence attached to it, and no model is involved in producing it.

What counts as attached. Evidence nested inside the card, and evidence joined to it by a connector in either direction. Direction is not a signal: people draw claim → quote as readily as quote → claim. Stance comes from the connector's label when the agent declared one, otherwise from the card's own supports/refutes pill.

How the score is built.

  • Independent sources drive it. Breadth has diminishing returns — one source is worth roughly half of the breadth term, three is worth about 88% of it. The fifth quote from the same forum adds very little; a second independent source adds a lot.
  • Independence is counted per community, not per domain. A subreddit is a source, so r/SaaS and r/startups count separately. Each Hacker News thread is a source. One app's review pool is a source, so reviews of two different apps are two sources even though both sit on one store domain.
  • Quote strength weighs in — weak, moderate or strong.
  • Old sources decay on a 180-day half-life, floored so nothing is discounted below half. An undated source is discounted slightly, not punished.
  • Source quality weighs in. A first-party or primary-voice page counts for more than an aggregator.
  • Refuting evidence subtracts. It is not filed away as a caveat.

What the labels mean.

LabelWhen
no evidence yetnothing attached
thinscored below 35
emerging35–69
solid70 or above
contestedrefuting evidence weighs at least as much as the supporting evidence

A claim can also carry a verdict from the experiments connected to it: unvalidated, testing, supported, refuted or contested. Refuted and contested claims are surfaced above the playbook checklist, because anything built on top of them needs review.

Connectors on the board

A connector is a line you draw — or the agent draws — between two cards. Drag from a card's edge to another card.

LabelMeaning
supportsThe source card is evidence for the target claim. Feeds confidence.
refutesThe source card is evidence against it. Subtracts from confidence.
derivationThe target was derived from the source. Carries a claim's verdict downstream to what was built on it.
(unlabelled)"These are related." Still counted for confidence when one end is an evidence card, using that card's own stance.

Connectors do double duty: they are the argument you can point at, and the signal the placement engine uses to decide where a card belongs. Labelling them is therefore worth the two seconds it costs.

The playbook: five stages

Five stages, sixteen steps that produce something, plus one "open the area" click per stage. Nothing runs on its own; each step names its objective, the command that starts it, and the condition that would finish it.

StageWhat it is forSteps
1 · ResearchMine what your market already saidCapture the idea · Mine the pain · Map the competition · Name who has the pain
2 · ValidateTurn a belief into something that can failReality-check the idea · State the riskiest assumption · Design the cheapest test · Record the verdict
3 · StrategyDecide from what survivedWrite the positioning · Set pricing · Define the identity · Plan go-to-market · Find funding you can apply to
4 · MarketingPut it in front of peopleSet the content strategy · Produce and publish posts
5 · AnalysisRead the result honestlyRead the results

Stage status is derived from the board on every render — which artifacts exist, how much evidence each carries, which experiments have verdicts. There is no second copy of your progress to drift out of date, and nothing to tick off by hand.

Stages arrive one at a time. Laying out all five up front produced four empty boxes for work you cannot sensibly start yet, so each next stage is added deliberately once the current one is underway.

Gates

A gate is what the playbook checks before it calls a step done.

StepGate
Mine the pain3 independent sources behind the pain map
Reality-check the idea2 independent sources
State the riskiest assumptionat least 1 source
Design the cheapest testmethod, participant source and metric all filled
Record the verdictan experiment linked, and its verdict in
Write the positioningsegment, benefit and statement filled
Plan go-to-marketmotion and channels filled
Find funding you can apply toprogramme, deadline and eligibility filled
Read the resultsan action named

Two properties matter more than the list:

  • Gates never block. You can move to the next stage with an unmet gate. You just do it knowingly, and the panel keeps saying what is missing.
  • Evidence one hop away counts. If the agent attached quotes to the four hypotheses hanging off your pain map rather than to the map itself, that is better research — each quote is bound to the specific claim it evidences — and the gate counts it, deduplicating sources so three quotes from one subreddit spread over three hypotheses still read as one source.

Agent modes

A mode decides which agent runs, and therefore which prompt and which tools it has. Three modes sit in the chat dropdown; the rest are slash commands, which are faster to reach when you want them and invisible when you do not.

In the dropdown

ModeWhat it does
WorkspaceWorks on the canvas: create, edit, organize. The default.
ResearchWeb and market intelligence.
Content creationCopy, media and documents.

Slash commands

CommandModeWhat it does
/validateValidateHypotheses, experiments, verdicts. Alias: /validation
/realityReality checkAn evidence-based verdict on the assumptions the idea rests on — including the founders who tried this and stopped. Aliases: /realitycheck, /check, /critique
/fundingFundingGrants, public programmes and investors, with real deadlines. Aliases: /grants, /investors
/strategyStrategyPositioning, pricing, go-to-market. Also owns the funding registries.
/marketingMarketingCampaigns and social posting.
/analysisAnalysisResults, insights, next steps.
/discoverDiscovery loopProbes each pain with a real post, reads the reaction, follows the evidence. Aliases: /discovery, /pains, /probe
/legislativeLegislativeRegulation and compliance. Prefers a large context window.

A few notes worth having:

  • The mode changes the tools, not just the tone. The funding step is /strategy for a reason: the funding registries live on that agent, and the default canvas agent has none of them — asked the same question it can only web-search and will write prose instead of cards.
  • /discover publishes. Probing a pain with a real post is an outward action, so it needs a connected account and passes through the approval gate.
  • Orchestrator modes delegate. Most stage modes are orchestrators: they spawn specialist sub-agents, each a real run with its own budget and journal, and merge the results. You see the delegation in the activity panel.

Research categories

Three commands narrow research to a category. They change the agent and the platforms it may reach — this is the only kind of command that touches the source set.

CommandReadsRough cost
/appApp Store and Google Play listings and reviewspaid store searches · ~10–20 tool calls
/saasReddit, Hacker News, pricing pages, GitHubmostly free sources · ~10–20 tool calls
/physical_productAmazon and eBay live prices and ratingspaid marketplace searches · ~10–20 tool calls

The category also filters the agent's tool manifest, which is a correctness fix as much as a cost one: product search is physical-product only and app search is app-only, because a SaaS run that searched a craft marketplace and an app idea researched with no store data at all both happened, and both cost real money.

How a run works

A run is a durable state machine, not a chat that has to stay open. Every model call and every side effect is a journaled step, so a run survives a provider timeout, a deploy, or you closing the tab.

What happens per turn: the agent is resolved from the registry (prompt, tools, language, identity), routed to a model, and then loops — the model either answers or asks for tools, the tools run as journaled steps, and their results feed the next turn.

What the agents can actually do. Tool names appear in the activity panel as the run works, so you can read what it touched:

GroupTools
Searchweb_search, google_search, tavily_search, read_result
Communitiesreddit_search, reddit_comments, hackernews_search, x_search, threads_search, instagram_research, linkedin_search
Marketapp_search, app_reviews, product_search, market_stats, cost_benchmarks, find_channels
Code & researchgithub_search
People & companiespeople_search, company_profile
Moneyfunding_search, national_funding_search, investor_search, startup_search
Board readsread_board, read_playbook, get_related_objects
Board writescreate_canvas_object, create_evidence, update_canvas_object, set_template_slot, set_todo_items, connect_canvas_objects, move_canvas_objects, set_canvas_parent, delete_canvas_objects
Memorysearch_document, read_document_chunk, search_conversations, read_conversation_chunk
Mediagenerate_image, generate_video

Guards that protect the work rather than the appearance of it:

  • An empty answer is not a success. A run that produced nothing says so.
  • A run whose every board write was rejected fails, whatever the model claims. The board is the deliverable.
  • Rejected writes are queued back to the model as a numbered work list, so a mistake is corrected rather than reported.
  • Deleting is a soft delete. Agent removals go to the board's trash, and a bulk removal asks you first.
  • Your undo is yours. Agent edits are a separate origin: Ctrl-Z never rewinds their work, and theirs never rewinds yours.

Research sources

Every source is a switch in Settings → Research. A source that is off is never called, whatever a run was asked for. Some switches cover several upstream sources — "Grants & funding" reaches the EU portal, grants.gov and CORDIS; "Official statistics" covers market size, technology adoption and research volume.

SourceCostWhat you get
RedditfreeCommunity discussions
Hacker NewsfreeFounder and practitioner threads
GitHubfreeCode and repositories
arXivfreeScientific papers
Official statisticsfreeMarket size, technology adoption and research volume by country
Channels & communitiesfreePodcasts and communities with audience sizes, so channels rank by reach
Startup directoriesfreeFunded competitors with batch, size and location
Grants & fundingfreeOpen public grants and calls you can apply for, matched to your country. Public money only, not investors
ThreadsfreeWellness, creator and lifestyle audiences. Needs a connected Threads account
InstagramfreeHashtag reach and engagement baselines. Needs a connected Instagram Business account
Web searchpaidGeneral web results, fast and broad
Web search (secondary)paidA second set of results alongside the main search
Page contentpaidFull text of a page — pricing pages, docs and specs
Marketplaces & app storespaidProduct listings, prices and reviews. Only used by category research, so it stays idle otherwise
LinkedIn profilespaidPeople and companies — roles, headcount and industry
LinkedIn postspaidPublic posts — B2B complaints and switching stories, with the author's job title
X (Twitter)paidPublic complaints and launch reactions with engagement counts
Agent web searchpaidLets the model search the web itself while it works

Three things worth knowing before you flip switches:

  • X is the most expensive source here by a wide margin, because it bills per post returned rather than per request. Turn it on when you need live reaction and engagement counts; leave it off otherwise.
  • Paid badges come from the backend, not from a hand-maintained list in the interface, so a source that changes its pricing cannot keep a stale badge.
  • Threads and Instagram show an "after deploy" badge where they are waiting on platform review. That is neither broken nor off by choice, and saying so beats a switch that silently achieves nothing.

Publishing connectors

A publishing connector is a social account you connect by OAuth so RawKit can post on your behalf and read back the numbers. Six platforms are implemented: LinkedIn, X, Reddit, Instagram, Threads and TikTok. No other platform can be connected.

How it works:

  1. You authorize on the platform. RawKit stores the access and refresh tokens encrypted, plus your account identifier, display name, avatar and the granted permissions. It never sees your password.
  2. Tokens are refreshed for you before they expire. Lifetimes differ per platform — hours on X, a day on TikTok, about sixty days on LinkedIn, Instagram and Threads — which is why refreshing is a background job rather than something you have to think about.
  3. Publishing is an outward action, and outward actions are gated. Nothing goes out as a side effect of a research run.
  4. Metrics come back onto the board. Views, likes, comments, replies, reposts, shares and scores — whatever the platform exposes — land next to the post card and the claim the post was testing. That is what makes an experiment's verdict a record rather than a feeling.
  5. You can disconnect at any time, which deletes the stored credentials. Instagram, Threads and TikTok also notify RawKit when you remove the integration from the platform's own settings; the connector is then revoked automatically.

The exact permissions each connector requests, what is stored, and what comes back are listed in section 8 of the Privacy Policy.

Not the same thing as a research source. Reading public LinkedIn posts, Threads or X for research is a source switch in Settings → Research. Posting as you needs a connector. Threads and Instagram research additionally needs a connected account, which is the one place the two overlap.

Configuration

Settings has four sections.

Profile

Your name, password, account — and your country, which is more consequential than it looks. See "Country and language" below.

AI models

Pick the model per job. OpenAI, Gemini, DeepSeek and MiniMax are wired.

ScopeUsed for
Default modelFallback for every AI task without a specific pick.
Canvas agentChat requests and board operations.
Research agentMarket, competitor and audience research.
Legislative agentRegulatory analysis — prefer a large context window.
Image generationThe generate_image tool.
Video generationThe generate_video tool. Its own scope, because clips cost an order of magnitude more than images and take minutes.
  • Every pick can carry a fallback, tried when the first choice is unreachable — a provider outage, a rate limit, an overloaded endpoint. Routing fails over by position, so a fallback is simply the next entry.
  • Some picks carry a "recommended" badge. Those come from running a real structured task against every model and scoring the output, not from taste. A model can be right for one mode and wrong for another, which is why the advice is per task — one model passes research and fails board-only work.
  • Only models that can actually do the job are offered. Every chat scope drives an agent with tools, so a model without function calling is not in the list rather than failing on turn one. Likewise an image model cannot be picked for video.
  • Routing is deterministic. It never spends a model call to decide which model to use: your preference first, then the platform default for that task, then price.

Research

The source switches above, plus research depth — the maximum number of agent turns per research run, from 1 up to the workspace cap.

Depth is the main cost lever, because every turn re-sends the conversation so far. Reaching the cap ends the run with a summary of what it found, never an error.

Usage

Your balance, what has been spent, and recharge.

  • Spend is broken out per model (tokens in and out) and per source (calls and dollars), so a workspace that spends more on searching than on thinking can see that.
  • A prepaid balance with a recharge threshold, topped up by card. Auto-recharge is optional.
  • When the balance runs low a run stops and hands back what it has — status stopped, with the partial result kept — rather than failing and discarding the work.

Cost and budgets

RawKit is metered. Every model call, paid search and media generation is priced at the provider's real price plus your plan's rate and drawn from the workspace's prepaid balance. The plan decides three things: the rate, which agents and research sources are available, and how many social accounts can be connected.

FreeCreativePro
Price$0 — pay as you go$9.99 / month$24.99 / month
Usage balance included$9.99 every month$24.99 every month
Usage rate over the provider's price+30 %+20 %+10 %
Top-ups$5–$500, card saved for optional auto-rechargesamesame
AgentsWorkspace, Research, Validate (incl. reality check), Strategy, Analysis, Legislative+ Content, Marketing+ Funding, discovery loop
Model tiersfast, primaryfast, primaryfast, primary, premium
Per-run spend cap$1$1$1
Concurrent runs1310
Image / video generationimagesimages and video
Research: community, web, app stores, marketplaces, statistics
Research: grant registries (grants.gov, CORDIS)
Research: social listening (X, Threads, Instagram, LinkedIn posts)
Research: investors (LinkedIn people and companies, startups database, model-native web search)
Connected social accountsup to 3: X, Instagram, Threads, TikTok, LinkedInunlimited, + Facebook, YouTube
Scheduler and publishing queue
Storage1 GB10 GB100 GB
Boards, sharing, multiplayer
  • The subscription comes back as balance. A paid plan's price is credited to the workspace's balance in full at the start of every billing period, so the fee is prepaid usage at a better rate plus the features — not a fee for access. Tax is added at checkout; the balance is always the pre-tax amount.
  • The plan lives on the workspace. Each workspace has its own plan, balance and billing. A user has no plan of their own; every member of a workspace works on its plan. A new workspace starts on Free.
  • A first research pass costs cents of model time. Every run reports its own tokens and dollars, per model and per data source, and the rate your plan applies to each model and source is shown in Settings.
  • Four independent budgets bound every run: spend (at most $1 per run, on every plan), turns, total steps, and wall-clock time. Any one of them trips a clean stop with a reason. A run that hits a wall says so instead of quietly truncating the work.
  • Locked features name the plan that has them. An agent, a source or a connector your plan does not include refuses with a message saying which plan does; nothing is silently downgraded. Sources your plan locks are marked as such in Settings.
  • Downgrade and cancellation. Plan changes, card updates, cancellation and invoices happen in the billing portal, reached from Settings. When a subscription ends the plan drops to Free at the end of the period. The balance you already have stays; further usage is charged at the Free rate; connected accounts stay connected but cannot post until you upgrade again.
  • Nothing runs unless you ask for it. There is no background crawl, no scheduled spend.
  • Repeated work is not re-charged. A retried run replays its journal instead of re-searching, and identical calls inside a run are deduplicated. A cached answer reports its real token counts at zero dollars, and the run says how many calls were cached — so a $0.00 line is never ambiguous.

Country and language

Set your country and national research runs in that country's language against localized search. This changes the results, not just the wording.

Measured: an English query for Austrian funding found almost none of the national agencies, while the German one found most of them. The words are the search key, not a courtesy translation — Förderung and Ausschreibung are what an Austrian agency calls its own money, dotace and výzva are the Czech pair, bando is what Italy calls a call. A dictionary does not produce these, so they are curated per country.

Two consequences worth knowing:

  • Country is per run, not only per workspace. A run that names its own country wins over the workspace setting, so you can research five markets from one board without changing your profile.
  • The layers are not equally machine-readable, and RawKit says which one a finding came from. Open EU calls come from the official portal with validated deadlines. Most national programmes publish no API at all, so those come from directed search against a registry of real agencies, promotional banks, angel networks and accelerators — and the deadline is read off the portal at request time rather than from a cache, because a stale deadline goes wrong invisibly.

Files and documents

Drop a PDF, image, spreadsheet or video onto the board and it becomes a card next to the work it belongs to. Text is extracted asynchronously, and the extracted text is what agents read — they can search a document and pull the passage they need rather than swallowing the whole file.

Conversations are searchable the same way, which is what makes a chat card worth keeping: a run can go back to what was said three sessions ago instead of you scrolling for it.

Deleting a project deletes its files with it. Deleting a file card leaves the underlying file intact; deleting the file removes it.

Working together, history and export

  • The canvas is real-time multiplayer. Two people editing the same board converge, including the same note's rich text, and an offline stretch merges on reconnect.
  • The agent is a participant, not a result dump. It appears in the presence bar as RawKit Agent while it works, and cards it wrote link back to the run that produced them.
  • Every agent edit is journaled with its origin, so the board can tell you which changes came from a run and which from a person.
  • Export: the board exports as a PNG, and every card is structured text you can select and copy — quotes with their urls included. There is no report generator.

When a run stops

A run that does not succeed tells you which wall it hit and what to do about it.

ReasonWhat it meansWhat to do
Spending limitThe run reached its dollar ceilingRaise it in Settings → Usage, or ask for a narrower scope
Step limitThe run reached its total step countNarrow the scope so it spends steps where they matter
Time limitThe run passed its wall-clock deadlineNarrow the scope, or split the work across two runs
Turn limitThe run used every model turn it hadAsk a more specific question
Out of fundsThe workspace balance ran outTop up in Settings → Usage
Stopped by youYou pressed StopAnything already on the board stays there
FailedAn errorThe reason is shown with the run

In every one of these cases the work already on the board survives. A stop is a soft landing by design: you keep what the agent produced up to that point.

What the beta does not include

Named here rather than discovered after signup:

  • Annual plans and team seats. Plans are monthly and per workspace; every member shares the workspace's plan. There is no annual price and no per-seat price.
  • A welcome credit. A new Free workspace starts with an empty balance and tops up from $5 before its first run.
  • A report export. PNG and copyable text, no generated document.
  • Private investor data. No free investor database exists that survives a coverage test, so investors come from a search that cites the page each name was found on — not from a database.
  • A few source connectors are waiting on platform review and carry an "after deploy" badge until they clear it.

Everything else in these docs is in the product today.

Glossary

TermMeaning
ArtifactA structured card with named fields — the comparable unit of work.
BoardOne infinite canvas; one idea or project.
Confidence0–100, computed from attached evidence. Never asserted by a model.
ConnectorA labelled line between two cards. supports and refutes feed confidence.
ContextThe cards you selected for a request to read.
EvidenceA verbatim quote with url, author, date, stance and strength.
GateA condition the playbook checks before calling a step done. Never blocking.
Independent sourceA distinct community or collection — a subreddit, a thread, one app's reviews, a publication.
ModeWhich agent handles a request, and therefore its tools and prompt.
OrchestratorAn agent that delegates to specialist sub-agents and merges their results.
RunOne durable, journaled, budgeted execution of one agent.
StageOne of the five playbook phases, each with its own frame on the board.
StanceWhether a quote supports or refutes the claim it is attached to.
VerdictA claim's status from the experiments linked to it: unvalidated, testing, supported, refuted or contested.
WorkspaceThe account-level container for settings, balance, connectors and boards.

Something wrong, missing or unclear in these docs? Tell us from inside the app — the feedback button is in the header. Docs are versioned, and this page says at the top which version you are reading.