AI can make a business experiment much cheaper. It can help you structure an idea, search for existing alternatives, draft neutral interview questions, and design a manual pilot before you build software.
What it cannot do is become the customer.
If a model says your idea sounds useful, generates ten enthusiastic fictional buyers, or predicts that people would pay £20, you have learned almost nothing about real demand. The useful role for AI is to improve the experiment, not manufacture the evidence.
In about an hour, a solo builder can turn one vague service idea into four practical assets: a falsifiable test card, a current alternatives map, a short customer-discovery script, and a manual pilot with an evidence ledger.
Capability check — August 7, 2026: OpenAI says ChatGPT Search is available across ChatGPT plans and can return current web information with source links. ChatGPT Projects can keep chats, files, and project instructions together and are available across free and paid subscriptions. Both are optional for this workflow: a normal text chat, browser search, and spreadsheet are enough. Usage limits and feature availability can vary by account.
1. Turn the idea into one falsifiable test card
Outcome: Replace “I think this could work” with one assumption you can actually try to disprove.
Best fit: Solo builders who have an idea but have not yet collected meaningful evidence from the target audience.
Inputs and tools: A one-sentence service idea, the audience you believe has the problem, and any general AI text assistant. A notes app is enough; a Project is useful only if you want to keep research and test records together.
Start by rewriting the idea in this form:
For [specific person], I will help with [specific recurring problem] by delivering [specific result].
Then ask the assistant to separate the assumptions underneath it. Typical categories are:
- the problem happens often enough to matter;
- you can reach the people who have it;
- they already spend time, effort, or money dealing with it;
- they will trust an outside person with the required inputs;
- you can deliver the result reliably;
- the result is valuable enough for them to take some real action.
Pick the assumption with the least evidence and the greatest ability to kill the idea. Do not test everything at once.
A useful interaction pattern is:
Do not improve or sell this idea. Convert it into a test card with: the assumption, why it must be true, the cheapest ethical test, observable evidence, what would disconfirm it, and fields where I must choose the success threshold. Mark assumptions as assumptions. Do not invent customer demand, conversion rates, prices, market size, or willingness to pay.
Before running the test, decide what result would make you continue, change direction, or stop. Strategyzer's Test Card uses the same basic discipline: state the hypothesis, define the experiment, identify what you will measure, and set the threshold before seeing the result.
Time and cost: About 10 minutes. Plain text is enough, so this can be done with an existing or free-tier AI account subject to its current limits.
Privacy, accuracy, and copyright limits: Do not feed the model confidential employer material, customer lists, private messages, credentials, or regulated data just to make the idea more specific. The model can help structure a test, but its opinion about whether the service is “promising” is not evidence.
2. Build a source-checked map of what people already do
Outcome: Understand the real alternatives to your service before assuming you have found an empty market.
Best fit: Service ideas where tools, pricing, competitors, or common workarounds can change quickly.
Inputs and tools: Current web search, the test card, and a simple table. ChatGPT Search can be convenient because it can return links alongside current information, but ordinary browser search works too.
Do not search only for direct competitors. A new service usually competes with several things:
- another paid service;
- software the customer already uses;
- a spreadsheet, template, or manual workaround;
- asking a colleague or friend;
- doing nothing because the problem is not painful enough.
Create a table with these columns:
| Alternative | Who it appears to serve | What it currently offers | Public price, if verified | Evidence URL and date | Question still unanswered |
|---|
Use official product or service pages for current features and prices when possible. Community discussions can reveal vocabulary and frustrations, but they are anecdotes, not a representative market survey.
A useful prompt is:
Search the current web for ways [audience] already handles [problem]. Include direct services, software, manual workarounds, and doing nothing. For every factual claim, provide the source and date. Prefer official pages for current features and pricing. If a current price or capability cannot be verified, write UNKNOWN. Do not infer demand from search visibility, review counts, or the existence of competitors.
The important question is not “Does a competitor exist?” It is:
What would make somebody switch from what they already do?
That gives you better discovery questions and may reveal that the useful offer is narrower than your first idea.
Time and cost: About 15 minutes for a first pass. OpenAI currently documents ChatGPT Search as available across plans, so there is a free route where access and usage limits permit. Browser search plus a spreadsheet is another low-cost route.
Privacy, accuracy, and copyright limits: Search results can be stale, incomplete, or geographically irrelevant. Re-open the source before relying on a feature or price. Summarise competitors in your own words rather than copying their landing-page text, screenshots, testimonials, or branded assets.
3. Draft discovery questions that ask about behaviour, not compliments
Outcome: A five-question conversation guide that can collect real evidence from people who actually experience the problem.
Best fit: Builders who can reach even a few relevant people through existing contacts, communities, customers, colleagues, or local groups.
Inputs and tools: The test card, alternatives map, and any text AI assistant.
AI is useful here because it can identify leading questions and rewrite them more neutrally. It should not answer the questions on behalf of fictional customers.
Weak discovery questions sound like this:
- Would you use an AI service that saved you time?
- Does this idea sound useful?
- Would you pay £20 for this?
Those questions encourage speculation and politeness.
Prefer questions about something that already happened:
- Tell me about the last time you dealt with this problem.
- What did you do first?
- What was annoying, slow, risky, or expensive about the current process?
- What did you try instead?
- What would have to happen for you to change the way you currently handle it?
Use AI to turn your specific idea into a short, neutral script:
Draft five customer-discovery questions about this problem. Ask about specific recent behaviour and the person's current workaround before any hypothetical purchase intent. Do not pitch the service inside the questions. Then draft a short outreach message that says I am researching the workflow, not promising an outcome. Flag every question that could bias the answer.
Keep the outreach small. You are not launching a campaign; you are trying to learn whether the problem behaves the way your test card assumes.
Time and cost: About 10–15 minutes to prepare the script and outreach message. A free text assistant is sufficient.
Privacy, accuracy, and copyright limits: Ask permission before recording or transcribing a conversation. Store only the information you need. If notes contain personal, confidential, or employer information, do not upload them to an AI service unless your agreement and account policy allow it. OpenAI's Data Controls can disable use of new personal-account chats for model improvement; Temporary Chat is not saved in history or used for training and is deleted from OpenAI systems after 30 days. Those product controls do not override an NDA, workplace rule, or customer promise.
4. Design a manual pilot before you build the product
Outcome: A tiny concierge-style test that lets a real participant experience the result before you automate the workflow.
Best fit: Reversible, low-risk services such as organising public research, preparing a non-regulated comparison, formatting content, setting up a simple workflow, producing a trip-planning draft, or turning supplied material into a structured checklist.
Avoid treating this method as a shortcut into legal, medical, financial, safety-critical, or other regulated work where qualified professional review may be required.
Inputs and tools: One consenting participant, the tools you already have, a spreadsheet or notes document, and an AI assistant for low-risk drafting or analysis.
Design the pilot in seven fields:
- Customer input: What exactly must the participant provide?
- Delivered result: What tangible output will they receive?
- AI-assisted work: Which parts can AI help draft, sort, compare, or format?
- Human review: What must you personally verify before delivery?
- Observable behaviour: What action would show that the result was useful enough to matter?
- Failure signal: What would show the service is too confusing, costly, slow, or untrusted?
- Decision rule: After the pilot, what would make you repeat, change, or stop?
Use this interaction pattern:
Design the smallest manual pilot for this service that can be delivered with existing tools. Separate: customer input, AI-assisted work, mandatory human review, delivered output, observable behaviour to record, failure signals, privacy risks, and stop/iterate criteria. Do not assume anyone will buy it, assign a success probability, or invent market demand.
Then create an evidence ledger:
| Date | Participant alias | What actually happened | Observable behaviour | Quote, only with permission | What this changes |
|---|
Synthetic cases are still useful for checking the workflow. You can ask an AI to invent awkward inputs and see whether your process breaks. Label those rows synthetic QA, however. They test the service design; they do not validate demand.
One real pilot is also weak evidence. It may teach you something useful, but it does not prove a market exists. The discipline is to record what happened without upgrading a small signal into a big claim.
Time and cost: About 20 minutes to design the pilot and ledger. The real recruitment and delivery will usually take longer and depends on the service. If you already have the required tools, the planning step can cost nothing beyond your existing access.
Privacy, accuracy, and copyright limits: Minimise personal data, verify every customer-facing factual claim, and do not use copyrighted source material beyond what you are allowed to process or reproduce. Never let an AI-generated deliverable imply that a licensed professional reviewed it when that did not happen.
A 60-minute starting session
| Activity | Time |
|---|---|
| Write the falsifiable test card | 10 minutes |
| Map current alternatives and evidence | 15 minutes |
| Draft discovery questions and outreach | 15 minutes |
| Design the manual pilot and evidence ledger | 20 minutes |
The hour does not validate the service. It prepares a cleaner experiment.
At the end, you should have:
- one assumption that could genuinely be disproved;
- a dated map of how people currently solve the problem;
- five neutral questions for real conversations;
- a manual pilot that can run before you build software;
- a place to record evidence without rewriting the story afterward.
The rule that keeps the experiment honest
A model-generated customer is useful for rehearsal, edge cases, and quality assurance. It is not a customer.
A model-generated market estimate is useful only as a hypothesis to verify. It is not market research.
A model saying “this sounds valuable” is not purchase intent.
The strongest use of AI is to make each experiment cheaper and clearer while preserving the distinction between generated output and observed human behaviour.
Build less before you know. Test more carefully. Record the failures as seriously as the positive signals.
AI should make the experiment cheaper—not turn the model into the customer.
Sources
Checked August 7, 2026:
- OpenAI Help Center — ChatGPT Search
- OpenAI Help Center — Projects in ChatGPT
- OpenAI Help Center — Data Controls and chat/file retention
- Strategyzer — Validate Your Ideas with the Test Card
- Strategyzer — How Strong Is Your Innovation Evidence?
- Strategyzer — Ways to Test Your Value Proposition and Business Model