Frame the objective
State the objective in one sentence, then set a strategy profile: platforms, build horizon, budget, privacy posture, and exclusions.
Every later step, hard gates included, reads from it.
Research and two independent AI evaluators pick the idea. The brand and screens are designed from that research. An AI agent builds the real app, and you ship it.
Runs on
The problem
Generating directions now takes seconds. Choosing the one that deserves a quarter of your team's time, and showing exactly why, is still the hard part.
The message that explained why an idea mattered now sits forty screens up a transcript no one will reopen.
An 8 out of 10 feels decisive until someone asks where it came from, and the room goes quiet.
Teams build the part they already understand and skip the one that could have ended the idea in a week.
The naming tool never saw the research, the design tool never saw the brand, and the app builder gets a paragraph. Each hop drops the evidence, the trade-offs, and the non-goals.
The path
Each stage reads what the last one made. Evidence picks the idea, the research names the brand, the brand styles the screens, and the agent builds against those screens. No tool starts from zero.
Stage 1 of 6 · Discover
Start from device and platform capabilities. Research runs store every source and extract atomic claims, each with a locator back to the passage.
In the project so far
Discover and decide
Language models generate and critique. Deterministic code owns the scores, the gates, and the provenance. You make the call with the full decomposition in view.
Choose from a living atlas of platform capabilities, each tracked for freshness and availability. Discovery combines them into opportunities worth researching.
Explore the capability atlasResearch runs store every source and extract atomic claims with locators. Generated text cites claim IDs, so nothing rests on assertion alone.
How evidence is storedUser, job, trigger, mechanism, wedge, business model, assumptions. Typed, versioned, forkable, and comparable side by side.
Inside a concept genomeOpenAI and Anthropic score every criterion independently. Deterministic code computes the weighted score and applies the hard gates. Disagreement stays visible rather than averaged away.
How evaluation worksAssumptions are ranked by importance, uncertainty, and the cost of being wrong. Every experiment carries kill criteria, and outcomes inform the next evaluation without rewriting the last.
Plan validationThe decision travels with the concept: a 26-section build spec, then the brand kit, confirmed screens, and specs, all imported into the Builder. Still exportable as a ZIP for any coding agent.
What the package containsThe workspace
The genome, the synthesized score, each criterion's decomposition, and the claims behind them share a single page, beside the actions that move a concept forward: promote it to validation, or generate its handoff.
Each criterion shows both evaluators, its weight, the median, and any gap worth a second look.
Every citation resolves to its stored source and the exact passage, with freshness marked.
Versions, forks, and decisions are recorded alongside the recommendation made at the time.
Brand Studio and Mockup Studio
The concept’s user, job, and evidence brief both studios. Brand Studio names it and builds the identity; Mockup Studio plans the app’s flows and designs each screen in that brand, ready for the Builder.
Candidates are checked for pronunciation and spelling in 12 languages, pre-screened against US and EU trademarks, and matched with live domain availability and prices. A first pass, not legal clearance.
The symbol is drafted with GPT Image 2 and traced to SVG. Lockups, monochrome, app icon, favicons, social avatar, and splash are then drawn exactly in code.
The palette is done only when every text and UI pair passes WCAG 2.2 AA, with type paired from an openly licensed catalog.
A storyboard of the app’s flows per device, each screen generated with your logo, palette, and type as references. Confirmed screens get editable specs.
Images use your plan’s image credits: a draft costs 6, a final 22, and the estimate is shown before every run. Assets drawn in code cost nothing.
The Builder
Your app opens in a live sandbox with the design package already in place. Ask for a change and watch an AI agent edit the code, check its work, and update the preview. Every request shows what it cost.
The brand kit, confirmed screens, and their specs land in docs/design, with icons installed, so the agent builds close to 1:1.
Pick a Claude or GPT-6 model at its list price, set the effort and a spend cap, and open any request’s cost breakdown.
Lint and typecheck run after every build turn. Every version keeps its diff and restores without rewriting history.
A session sleeps after 3 minutes without activity and wakes when you return, so a forgotten tab stops using build minutes.
Start from a concept, a prompt, or a template.
The monorepo pairs an Express API with the apps you choose: web, admin, marketing, an Expo mobile app, and a worker.
Step by step
The path in six steps. Each one produces durable, linked records that the next one reads: nothing to dig out of a transcript, no score without a source, no tool starting from zero.
State the objective in one sentence, then set a strategy profile: platforms, build horizon, budget, privacy posture, and exclusions.
Every later step, hard gates included, reads from it.
Choose capabilities from the atlas. Research runs collect sources and extract cited claims, and opportunities become structured concept genomes.
Retrieved pages are treated as data, never as instructions.
OpenAI and Anthropic evaluators score each concept independently, and code computes the result. The riskiest assumptions get experiments with kill criteria.
Outcomes inform a new evaluation and never rewrite the previous one.
Brand Studio names the concept and builds its identity from the research. Mockup Studio plans the flows and designs each screen in that brand.
Names are pre-screened, palettes pass WCAG AA, confirmed screens carry specs.
The Builder imports the design package into a live sandbox, and an AI agent builds the app while you watch the preview and steer.
Lint and typecheck after every build turn, and every request’s cost shown.
Open a pull request on your GitHub repository, download the code as a ZIP, or deploy a shareable preview with its own database.
Every publish is recorded with its report and repository ZIP.
The difference
Most AI tools return fluent prose you cannot audit, and each one starts from zero. LiteSurface keeps reasoning and evidence apart and carries them from the first claim to the running app.
Where ideas live
Usual way: In a chat transcript you have to scroll back through to find again.
What backs a score
Usual way: A confident number with nothing behind it.
Who judges
Usual way: One model’s opinion, delivered as fact.
Deal-breakers
Usual way: A persuasive pitch can hide a fatal constraint.
Risk
Usual way: Build first, find out later.
Handoff
Usual way: A document that loses its context on the way.
Design
Usual way: A naming tool, a logo tool, and a design tool, none of which saw the research.
Build
Usual way: A ZIP for someone else to start over from.
Where ideas live
LiteSurface way: Durable, linked objects: capabilities, sources, claims, concepts, and decisions.
What backs a score
LiteSurface way: Claims cite stored sources with locators, and evidence coverage is measured.
Who judges
LiteSurface way: Two independent evaluators, OpenAI and Anthropic. Code synthesizes; disagreement stays visible.
Deal-breakers
LiteSurface way: Hard gates (build horizon, legal risk, acquisition path) fail a concept regardless of score.
Risk
LiteSurface way: The riskiest assumptions are tested first, with kill criteria set before results arrive.
Handoff
LiteSurface way: A 26-section build spec pinned to the concept version, read by the Builder.
Design
LiteSurface way: A brand and screens designed from the research, with names pre-screened and a palette that passes WCAG AA.
Build
LiteSurface way: A running app built from your research and screens, with every request’s cost shown.
Take the method apart.
Weights, gates, medians, and coverage: every rule behind a recommendation.
Weighing it against something specific? LiteSurface vs. AI app builders, LiteSurface vs. spreadsheets, LiteSurface vs. ChatGPT or Claude chat, or LiteSurface vs. consultants.
Try the scorecard
Adjust each criterion the way an evaluator might. The synthesized score, band, and recommendation follow the same rules as the product, in miniature.
OpenAI and Anthropic score each criterion independently. Code takes the median and applies the weights; no model writes the final number.
A failed hard gate, such as a prototype that overruns your build horizon or an unacceptable legal risk, overrides any total.
Confidence, disagreement, and evidence coverage sit beside the score and never inflate it. Thin evidence caps the recommendation and sends the concept back for research.
Interactive scorecard
Offline voice notes for field consultants
Synthesized score 72 out of 100, Strong. Recommendation: Validate.
A simplified model. The product’s default scorecard weighs nine criteria, takes per-criterion medians from two evaluators, reports confidence and disagreement, and applies six gates. Evidence coverage acts as a soft gate that caps the recommendation and triggers an evidence-gap pass. At 85 and above, the product recommends building or validating depending on risk and evidence.
Ready for the full version?
Two independent evaluators, nine criteria, six gates, every score cited.
Pricing
Solo and Team cost nothing while we build in the open. Every plan includes build minutes for the Builder and image credits for the studios; model calls run on the provider keys you connect and are metered in the usage ledger. Prices shown for later are indicative.
Free
during early access
One workspace and unlimited projects for a founder or a single product lead.
Free
during early access
Then $49 per seat / month, billed annually (indicative)
Shared workspaces, roles, and the admin console for product teams.
Custom
annual agreement
Dedicated deployment, custom scorecards, managed keys, and procurement support.
Bring your own provider keys; budgets cap what any workspace can spend.
Compare plansQuestions
What the system does, how it reaches a recommendation, and what it costs during early access.
Prefer the long version?
Each stage of the pipeline, and the guarantee that comes with it.
How it worksA system that takes a product idea to production in one project. You choose device and platform capabilities from the atlas; the engine researches the web for evidence, turns opportunities into structured concept genomes, scores them with two independent evaluators, and plans validation. Brand Studio and Mockup Studio then design the brand and the screens from that research, and the Builder’s AI agent builds the real app against them, ready to ship.
Every meaningful result is a durable, linked object, so each stage reads what the last one made instead of starting from zero.
Two independent evaluators, one from OpenAI and one from Anthropic, score each criterion from 0 to 100 against a weighted scorecard. Code, not a model, takes the median per criterion and computes the weighted total, confidence, disagreement, and evidence coverage.
Hard gates, such as a prototype that overruns your build horizon or an unacceptable legal risk, override the number. Thin evidence caps the recommendation and triggers an evidence-gap pass that researches what is missing, then re-evaluates. Bands: 85 and up is exceptional, 72 strong, 58 hold, and anything lower is weak.
Real, runnable web apps in a live sandbox, in Next.js, Vite + React, or Astro. Start from a blank app, a starter (a SaaS dashboard, CRM, internal tool, task board, landing page, or blog), a prompt, or a concept. The full-stack monorepo template pairs a required Express 5 API with the apps you choose: a web app, an admin app, a marketing site, an Expo mobile app, and a worker, with auth, database, and email options and optional payments, storage, analytics, and error tracking.
The agent works in Build, Plan, or Chat mode. In Build mode it edits the code, saves a version, and runs lint and typecheck, while you watch the preview, point it at an element, or restore any earlier version.
Brand Studio screens name candidates (pronunciation and spelling in 12 languages, a US and EU trademark pre-screen, live domain availability and prices), builds a palette that must pass WCAG 2.2 AA, pairs type from an openly licensed catalog, and drafts a logo symbol with GPT Image 2 that is traced to SVG. Lockups, icons, favicons, and the other derived assets are drawn exactly in code. Mockup Studio generates each screen with GPT Image 2, using your logo, palette, and type as references.
Under our Terms you own your content and the output generated for you, and we assign any rights we have in that output to you. Outputs may not be unique, and the name screen is a first pass, not legal clearance: have counsel review a name before you adopt it.
Yes. The Builder’s publish menu downloads the repository as a ZIP on any plan, or commits to a branch and opens a pull request on a connected GitHub repository. On Team and Enterprise you can also deploy a shareable preview with its own Postgres database; claim it and its hosting and database move into your own accounts.
Previews of the full-stack monorepo are not deployable yet; push those to GitHub or download the ZIP.
Every request shows its cost, with a breakdown by model and token type. You choose the model (Claude or GPT-6 models, each at its list price), the effort, and a spend cap per request, which defaults to $2 and can go up to $10. Workspace budgets and the usage ledger cover all model spend.
Sandbox time counts against your plan’s build minutes (60, 600, or 3,000 a month), and a session sleeps after 3 minutes without activity, so an idle tab stops using them.
OpenAI and Anthropic evaluate concepts, research runs on Tavily or Brave, the studios generate images with OpenAI’s GPT Image 2, and the Builder’s agent runs on your choice of Anthropic Claude or OpenAI GPT-6 models. If one evaluator’s provider is unavailable, evaluation falls back to the other and continues in single-evaluator mode, flagged on the card.
Each workspace can connect its own keys. During early access, calls without a workspace key can use keys the deployment manages; after early access that is an Enterprise option. Budgets and the usage ledger show what each project and run costs.
Projects are the unit of ownership. Transfer a project and every concept, source, evaluation, experiment, and artifact moves with it.
Only the fields a task needs are sent to a model. Fields classified as sensitive are never sent to model providers, and provider keys stay server-side, never stored in plaintext. In the Builder, the agent sees the names of your environment variables, never their values.
Yes. Explore a sample project before you run your own: concepts, both evaluators’ scores, gates, citations, and a handoff, all in the real product. You can also read one complete sample evaluation on this site without signing up.
Self-hosting is available for teams that need it, including a fixture mode that replaces model calls with deterministic outputs.
A 26-section implementation spec covering product, UX, engineering, AI, quality, execution, and sources, pinned to the concept version it was built from. Then the published brand kit (DESIGN.md guidelines, tokens, logo files, and fonts) and the confirmed screens with their specs, imported into the app’s docs/design folder, with favicons, app icons, and the social image installed.
The spec still exports as a ZIP with footnoted citations, so any coding agent can use it.
Solo and Team are free during early access. After that, Team is indicatively $49 per seat per month, billed annually, and Enterprise is a custom annual agreement.
Each plan includes build minutes and image credits every month: Solo 60 and 400, Team 600 and 4,000, Enterprise 3,000 and 20,000. Model calls run on the provider keys you connect, capped by workspace budgets. Prices may change before general availability.
Founders, product leads, and product teams who want to go from a question to a shipped product with the evidence attached: a shortlist they can defend, a brand and screens built on the research, and an app built against them.
Explore a sample project before you run your own, then take your idea from evidence to a brand, screens, and a running app, in one project.