You have made one. It is on the kitchen table where the light is good, and it is genuinely nice, and you have picked it up four times in the last hour to look at it again. Then the other thought arrives, the one that does not go away: how many should I make?
The honest answer is that you do not know yet. You have a feeling. Feelings have talked a lot of people into a hundred units of something.
If there is a box in the garage from the last idea — the one you still cannot bring yourself to throw out because it cost real money and represents four weekends — you already know how this goes. That box is not evidence that you have bad instincts. It is evidence that you made a decision before you had anything to decide with.
This post is about the three weeks in between. Validating a handmade product idea means selling a small, constraint-sized test batch at your real price, to strangers, across at least two selling occasions — with the numbers that will decide it written down before anything sells. Not a business plan, not a survey, not a focus group. Three weeks, about a quarter of the units, and a filled-in card that makes the decision for you.
Why 100 is the wrong first number
A hundred is not a considered quantity. It is the number that arrives when a supplier's minimum is a hundred, or when round numbers feel serious, or when someone at a market says you should make loads of these and you believe them because you wanted to.
Count what a hundred actually costs. Materials, obviously. But also several weeks of the only production capacity you have, which is the same capacity your existing sellers need. The cash, which is now sitting on a shelf instead of in the account. The storage. And the least visible one: commitment. Once a hundred units exist, you will defend them. You will discount them, bundle them, take them to markets that are not worth the booth fee, and generally spend a year proving you were right instead of finding out.
Closing up is common enough to be ordinary: of the private-sector establishments that opened in the year ending March 2024, 77.9% were still open a year later, and of those that opened in the year ending March 2020, 51.4% survived to March 2025 (BLS Business Employment Dynamics, Table 7 (opens in new tab)). Those figures count businesses, not product lines, so they say nothing direct about your mugs. But a small business is largely a stack of product decisions, and the quiet endings are rarely dramatic. They look like a garage.
The four questions a test is actually for
It is tempting to test one thing and call it validation: did anyone buy it? That is the easiest question to pass and the least useful one to stop at. A real test is trying to answer four, in this order:
- Will a stranger buy it at all? Friends buy the maker. You need people with no relationship to you and no social cost for walking past.
- At what price? Not "would you pay $30 for this" — nobody knows what they would pay. At the price you charged, what happened.
- Will they buy it again? This is the question that separates a product line from a novelty, and it is the one a single market day cannot answer.
- Can you make it at that cost, at that speed, repeatedly? The first one was made lovingly over an evening. The eightieth has to be made on a Tuesday when you are tired.
Question 4 is the one a test can most easily get a false pass on, because the first unit is always cheaper and better than the average unit — it was made slowly, by someone paying attention.
Building the test batch
Size it from a constraint, not from a hope. One kiln load. One cook. One cut of fabric. One resin pour. Physical constraints make good batch sizes because they are the same quantity you would have to repeat, so the test measures the real production cycle rather than a one-off heroic effort.
If nothing constrains you, the arithmetic sets sensible bounds. With twelve units, a single sale moves sell-through by 8.3 percentage points, so one chatty buyer can flip your verdict — that is noise, not a signal. With twenty-four, a unit is 4.2 points, and the shape of the result starts to mean something. Beyond about thirty you are no longer testing; you have started producing and are looking for permission.
Price it where you want to live. Test at the price you would actually charge, not a "just to get feedback" discount. A discounted test tells you exactly one thing, which is how people feel about a discount. If the price is wrong you want to find that out now, and you cannot find it out at a price you will never charge again.
Sell it to strangers, twice. One outing answers question 1. Only the second one begins to answer question 3. Three to five weeks is usually enough for two or three markets, or a listing that has been live long enough to stop being novel to your own followers.
Write the thresholds down before anything sells. This is the whole trick, and it is the step most easily skipped. Three lines, on paper, before the batch leaves the house: the number at which you make more, the number at which you stop, and the signal that means you were underpriced. Decide them while you are still capable of being disappointed. After the market you will be either elated or crushed, and neither of those states makes good decisions.
The scorecard, torn down
Nadia throws stoneware in a converted garage, six hours a week, around a job. She had made one mug with a thumb rest that people kept picking up. Instead of making a hundred, she made a kiln load and filled in a card. (Nadia and the two makers after her are composites, built to show the shape of a decision rather than to report anyone's actual quarter.)
| # | Line | What she wrote down |
|---|---|---|
| (1) | Batch size and why | 24 — one glaze load |
| (2) | Test price | $28 |
| (3) | Thresholds, set 9 days before the market | Sell 15+ → make more · Sell under 8 → stop · Sell out before 1pm → raise the price |
| (4) | All-in cost per unit | $12.61 |
| (5) | Units sold / offered | 24 of 24 |
| (6) | Sold out at | 11:40am, of a 9am–3pm market |
| (7) | Asked after they were gone | 10 |
| (8) | Verdict against (3) | Underpriced — same batch size, higher price, run it again |
(1) The batch size came from the kiln, not from ambition. Twenty-four is what fits. It is also what she can fire in one cycle, which means the test measured a repeatable production run rather than a special effort.
(2) $28 was the price she wanted, arrived at before the test rather than during it. She did not put out a "market special" sign.
(3) The thresholds were written nine days early, which is the single most important row on the card. Note that there are three of them, not two. Most people write a pass line and a fail line and have no plan for the third outcome, which is the one that actually happened.
(4) The all-in cost is not a guess. Clay $0.68, glaze $0.41, her share of two kiln firings $0.64, packaging $1.35, and 22 minutes of hands-on work at $26 an hour — $9.53. Total $12.61. If you want to build the same figure for your own product, the Product Pricing Calculator will walk the same lines in a browser. In Ardent Seller you would build the test batch as a recipe once; it pulls the material costs from what you actually paid your suppliers and re-costs the mug on its own when clay goes up, which matters more than it sounds like it does — the whole verdict below hangs off this one number.
(5) and (6) look like a triumph and are not. Twenty-four of twenty-four is a 100% sell-through, and it happened in 2 hours 40 minutes of a six-hour market. She spent the remaining three hours and twenty minutes standing behind a table with nothing to sell, having paid $65 for the booth.
(7) is the row most people skip — call it the asked-after-gone count — and it is worth more than most of the others. Ten more people asked for a mug after they were gone. That is not a rounding error; it is nearly 42% again on top of the batch, and it converts a fuzzy feeling into a line you can act on.
(8) The verdict was written by row 3, not by Nadia. That is the point of writing thresholds down. Her contribution on the day was 24 × ($28 − $12.61) = $369.36, less the booth, less card fees — a little under $290. But her $12.61 covers making the mugs, not the six hours plus setup she spent selling them. Charge those at the same $26 an hour and the day cleared roughly $130. The mugs were not the problem. The price was.
Three makers, three verdicts
Nadia: sold out, and that was the bad news
A fast sellout feels like the best possible result and is usually a pricing signal. If everything goes in the first third of the day and ten more people ask afterwards, the market was telling her the mug was worth more than $28 — and she had no way to find out how much more, because she had nothing left to test with.
Her second run was the same 24 mugs at $36. Twenty sold, four came home, and four came home deliberately: having stock at 3pm is what lets you learn where the ceiling is. The four went into an online listing that week. The verdict was not "make a hundred." It was "make twenty-four again, charge more, and keep some."
Theo: everyone bought one, almost nobody bought two
Theo cooks hot sauce. His new one was a ghost-pepper thing with a skull on the label, and it did well: 34 of 40 bottles across three markets in five weeks, an 85% sell-through that would pass almost any threshold anyone would write.
But he had also written down a fourth line, which was to count returning buyers. He asked at the till. Market one: 14 sold, all new faces. Market two: 12 sold, one returning. Market three: 8 sold, one returning. Two repeat buyers out of 34. Over the same three markets his garlic-dill — a sauce nobody has ever described as exciting — ran closer to one in three.
That is not a failed product. It is a product that turns out to be a gift and a dare rather than something that lives in a fridge door. So the verdict was neither yes nor no: keep it, make it in small batches, put it at the front of the table where it pulls people in, and do not build the next year of the business around it. He would never have found that out from sell-through alone, and a hundred bottles of it would have been a slow, confusing year.
Ros: the test that never became a batch
Ros designs embroidered patches and wanted to make an enamel pin. Enamel pins do not permit a small test batch: the factory wanted a mold fee plus a minimum of 100 pieces, about $295 all in before a single one had been shown to anybody. There is no 24-unit version of that decision.
So she moved the test upstream. She finished the artwork, posted it, said plainly that she would order the mold if 40 people said they wanted one, and counted. No money changed hands — she was measuring intent, not taking orders. Three weeks: about 1,140 people saw it, 61 liked it, 9 said they wanted one.
Nine. Against a threshold of 40 that she had written down before posting, which is the only reason nine was allowed to mean anything. Had she picked the number afterwards, nine enthusiastic comments would have felt like momentum.
The test cost her an evening of drawing and three weeks of patience. It saved $295 and a box of 100 pins. If you do decide to take money on a test like this rather than just names, the mechanics stop being a marketing question and become a legal one — the FTC's Mail, Internet, or Telephone Order Merchandise Rule (16 CFR Part 435 (opens in new tab)) governs what you must ship by when and what you owe if you cannot. The guide to taking preorders for handmade products covers how preorders, waitlists and backorders actually differ.
What the three tests had in common
None of them was expensive, and none of them produced a yes-or-no.
Every one of the three set the deciding number before the test ran. Nadia had three thresholds nine days out; Theo decided to count returning faces before market one; Ros published her number publicly, which made it very hard to move. This is the mechanism, and it works because it moves the judgment call to the only moment when you are neutral about the answer.
Every one measured something other than enthusiasm. Sell-through, minutes to sell out, asked-after-gone, returning buyers, committed names against a stated target. None of those is did people seem to like it, which is the metric a craft fair generates in abundance and which almost never changes anyone's decision.
And each test cost a small fraction of the commitment it was standing in for. Nadia risked 24 mugs against a season of production. Theo risked 40 bottles. Ros risked nothing but three weeks, against $295 and a drawer.
Two of the three verdicts were "yes, but differently" — a different price, a different role in the range. Only one was a no, and it was the cheapest one to receive.
Reading your own verdict
Four results, four different next moves.
Sold out fast, and people asked afterwards. You were underpriced. Do not scale the batch — scale the price, keep the batch size, and deliberately plan to have stock left at the end so you can see where the ceiling is.
Sold most of it, at a pace, and some people came back. This is the boring result that actually means "make more." Scale to your next physical constraint, not to a hundred: the next kiln load up, the next cook size, the next fabric roll.
Sold most of it, and nobody came back. You have a novelty, and novelties are useful. Keep it in the range, keep the batches small, and let it do the job of stopping people at the table. Do not let it become the plan.
Barely sold. Work out whether anyone stopped at all. If people walked past without slowing, the problem is the venue or the audience and you should retest somewhere else before you touch the product. If people picked it up, turned it over, and put it down, you have your answer and it is a cheap one.
Whichever you get, run the winner against what you already sell before you commit the production time. In Ardent Seller that comparison lives in the sales reports — sell-through and margin per product, side by side — and it is a genuinely uncomfortable screen, because a new product that everybody complimented will sometimes sit below the boring thing you have made for four years. Better to find that out over 24 units than over a season.
The box in the garage was never about instinct. It was about deciding first and finding out afterwards. Make one kiln load, write three lines on a card, and let the card decide.
Ready to know what your test batch actually costs before you make it? Start free and build the batch as a recipe — materials, packaging and your own time — so the cost-per-unit on your scorecard is a figure rather than a guess.
Related reading
- How to Take Preorders for Handmade Products — if your test involves taking money before the thing exists, this is the difference between a preorder, a waitlist and a backorder, and what each one obliges you to do.
- How Many Copies Should Your First Board Game Print Run Be? — the same commitment problem where the factory minimum is thousands, not one kiln load, and the argument for sizing nearer the MOQ than the internet recommends.
- Should You Launch a Handmade Subscription Box? — a four-number go/no-go framework for the case where the product is a recurring commitment rather than a single item.
- Recipe Costing 101 — how to build the all-in cost-per-unit figure that row 4 of the scorecard depends on, across all five cost categories.
Free resources
Free companion downloads if you want to run a test batch this month:
- Craft Show Prep & Profit Tracker — a place to record units offered, units sold, booth costs and the asked-after-gone tally on the day, while you still remember them.
- Should I Raise My Prices? Decision Tool — for the sellout verdict, when the test says you were underpriced and you need to decide how far to move.
- Small-Batch Production Planning Playbook — for when the test passes and the next batch has to fit around everything you already make.
This article is provided for educational purposes only and does not constitute legal, regulatory, financial, or tax advice. Costs, pricing examples, and federal preorder and mail-order rules vary by jurisdiction and change frequently. Consult a qualified accountant, small-business advisor, or attorney before making financial or compliance decisions based on this content.
