Skip to content
Growth · 17 min read

How to Validate a Handmade Product Before You Make 100 of Them

A test-batch protocol for makers who have made exactly one of something and do not know what to do next: how big the batch should be, what to charge, which numbers to write down, and why the thresholds have to be set before you sell anything.

Four bars of handmade soap in different colors, each wrapped in kraft paper and tied with twine, lined up on a weathered wooden bench beside sprigs of dried lavender

You have made one. It is on the kitchen table where the light is good, and it is genuinely nice, and you have picked it up four times in the last hour to look at it again. Then the other thought arrives, the one that does not go away: how many should I make?

The honest answer is that you do not know yet. You have a feeling. Feelings have talked a lot of people into a hundred units of something.

If there is a box in the garage from the last idea — the one you still cannot bring yourself to throw out because it cost real money and represents four weekends — you already know how this goes. That box is not evidence that you have bad instincts. It is evidence that you made a decision before you had anything to decide with.

This post is about the three weeks in between. Validating a handmade product idea means selling a small, constraint-sized test batch at your real price, to strangers, across at least two selling occasions — with the numbers that will decide it written down before anything sells. Not a business plan, not a survey, not a focus group. Three weeks, about a quarter of the units, and a filled-in card that makes the decision for you.

Why 100 is the wrong first number

A hundred is not a considered quantity. It is the number that arrives when a supplier's minimum is a hundred, or when round numbers feel serious, or when someone at a market says you should make loads of these and you believe them because you wanted to.

Count what a hundred actually costs. Materials, obviously. But also several weeks of the only production capacity you have, which is the same capacity your existing sellers need. The cash, which is now sitting on a shelf instead of in the account. The storage. And the least visible one: commitment. Once a hundred units exist, you will defend them. You will discount them, bundle them, take them to markets that are not worth the booth fee, and generally spend a year proving you were right instead of finding out.

Closing up is common enough to be ordinary: of the private-sector establishments that opened in the year ending March 2024, 77.9% were still open a year later, and of those that opened in the year ending March 2020, 51.4% survived to March 2025 (BLS Business Employment Dynamics, Table 7 (opens in new tab)). Those figures count businesses, not product lines, so they say nothing direct about your mugs. But a small business is largely a stack of product decisions, and the quiet endings are rarely dramatic. They look like a garage.

The four questions a test is actually for

It is tempting to test one thing and call it validation: did anyone buy it? That is the easiest question to pass and the least useful one to stop at. A real test is trying to answer four, in this order:

  1. Will a stranger buy it at all? Friends buy the maker. You need people with no relationship to you and no social cost for walking past.
  2. At what price? Not "would you pay $30 for this" — nobody knows what they would pay. At the price you charged, what happened.
  3. Will they buy it again? This is the question that separates a product line from a novelty, and it is the one a single market day cannot answer.
  4. Can you make it at that cost, at that speed, repeatedly? The first one was made lovingly over an evening. The eightieth has to be made on a Tuesday when you are tired.

Question 4 is the one a test can most easily get a false pass on, because the first unit is always cheaper and better than the average unit — it was made slowly, by someone paying attention.

Building the test batch

Size it from a constraint, not from a hope. One kiln load. One cook. One cut of fabric. One resin pour. Physical constraints make good batch sizes because they are the same quantity you would have to repeat, so the test measures the real production cycle rather than a one-off heroic effort.

If nothing constrains you, the arithmetic sets sensible bounds. With twelve units, a single sale moves sell-through by 8.3 percentage points, so one chatty buyer can flip your verdict — that is noise, not a signal. With twenty-four, a unit is 4.2 points, and the shape of the result starts to mean something. Beyond about thirty you are no longer testing; you have started producing and are looking for permission.

Price it where you want to live. Test at the price you would actually charge, not a "just to get feedback" discount. A discounted test tells you exactly one thing, which is how people feel about a discount. If the price is wrong you want to find that out now, and you cannot find it out at a price you will never charge again.

Sell it to strangers, twice. One outing answers question 1. Only the second one begins to answer question 3. Three to five weeks is usually enough for two or three markets, or a listing that has been live long enough to stop being novel to your own followers.

Write the thresholds down before anything sells. This is the whole trick, and it is the step most easily skipped. Three lines, on paper, before the batch leaves the house: the number at which you make more, the number at which you stop, and the signal that means you were underpriced. Decide them while you are still capable of being disappointed. After the market you will be either elated or crushed, and neither of those states makes good decisions.

The scorecard, torn down

Nadia throws stoneware in a converted garage, six hours a week, around a job. She had made one mug with a thumb rest that people kept picking up. Instead of making a hundred, she made a kiln load and filled in a card. (Nadia and the two makers after her are composites, built to show the shape of a decision rather than to report anyone's actual quarter.)

Nadia's completed test-batch scorecard, 24 stoneware mugs
# Line What she wrote down
(1) Batch size and why 24 — one glaze load
(2) Test price $28
(3) Thresholds, set 9 days before the market Sell 15+ → make more · Sell under 8 → stop · Sell out before 1pm → raise the price
(4) All-in cost per unit $12.61
(5) Units sold / offered 24 of 24
(6) Sold out at 11:40am, of a 9am–3pm market
(7) Asked after they were gone 10
(8) Verdict against (3) Underpriced — same batch size, higher price, run it again

(1) The batch size came from the kiln, not from ambition. Twenty-four is what fits. It is also what she can fire in one cycle, which means the test measured a repeatable production run rather than a special effort.

(2) $28 was the price she wanted, arrived at before the test rather than during it. She did not put out a "market special" sign.

(3) The thresholds were written nine days early, which is the single most important row on the card. Note that there are three of them, not two. Most people write a pass line and a fail line and have no plan for the third outcome, which is the one that actually happened.

(4) The all-in cost is not a guess. Clay $0.68, glaze $0.41, her share of two kiln firings $0.64, packaging $1.35, and 22 minutes of hands-on work at $26 an hour — $9.53. Total $12.61. If you want to build the same figure for your own product, the Product Pricing Calculator will walk the same lines in a browser. In Ardent Seller you would build the test batch as a recipe once; it pulls the material costs from what you actually paid your suppliers and re-costs the mug on its own when clay goes up, which matters more than it sounds like it does — the whole verdict below hangs off this one number.

(5) and (6) look like a triumph and are not. Twenty-four of twenty-four is a 100% sell-through, and it happened in 2 hours 40 minutes of a six-hour market. She spent the remaining three hours and twenty minutes standing behind a table with nothing to sell, having paid $65 for the booth.

(7) is the row most people skip — call it the asked-after-gone count — and it is worth more than most of the others. Ten more people asked for a mug after they were gone. That is not a rounding error; it is nearly 42% again on top of the batch, and it converts a fuzzy feeling into a line you can act on.

(8) The verdict was written by row 3, not by Nadia. That is the point of writing thresholds down. Her contribution on the day was 24 × ($28 − $12.61) = $369.36, less the booth, less card fees — a little under $290. But her $12.61 covers making the mugs, not the six hours plus setup she spent selling them. Charge those at the same $26 an hour and the day cleared roughly $130. The mugs were not the problem. The price was.

Three makers, three verdicts

Nadia: sold out, and that was the bad news

A fast sellout feels like the best possible result and is usually a pricing signal. If everything goes in the first third of the day and ten more people ask afterwards, the market was telling her the mug was worth more than $28 — and she had no way to find out how much more, because she had nothing left to test with.

Her second run was the same 24 mugs at $36. Twenty sold, four came home, and four came home deliberately: having stock at 3pm is what lets you learn where the ceiling is. The four went into an online listing that week. The verdict was not "make a hundred." It was "make twenty-four again, charge more, and keep some."

Theo: everyone bought one, almost nobody bought two

Theo cooks hot sauce. His new one was a ghost-pepper thing with a skull on the label, and it did well: 34 of 40 bottles across three markets in five weeks, an 85% sell-through that would pass almost any threshold anyone would write.

But he had also written down a fourth line, which was to count returning buyers. He asked at the till. Market one: 14 sold, all new faces. Market two: 12 sold, one returning. Market three: 8 sold, one returning. Two repeat buyers out of 34. Over the same three markets his garlic-dill — a sauce nobody has ever described as exciting — ran closer to one in three.

That is not a failed product. It is a product that turns out to be a gift and a dare rather than something that lives in a fridge door. So the verdict was neither yes nor no: keep it, make it in small batches, put it at the front of the table where it pulls people in, and do not build the next year of the business around it. He would never have found that out from sell-through alone, and a hundred bottles of it would have been a slow, confusing year.

Ros: the test that never became a batch

Ros designs embroidered patches and wanted to make an enamel pin. Enamel pins do not permit a small test batch: the factory wanted a mold fee plus a minimum of 100 pieces, about $295 all in before a single one had been shown to anybody. There is no 24-unit version of that decision.

So she moved the test upstream. She finished the artwork, posted it, said plainly that she would order the mold if 40 people said they wanted one, and counted. No money changed hands — she was measuring intent, not taking orders. Three weeks: about 1,140 people saw it, 61 liked it, 9 said they wanted one.

Nine. Against a threshold of 40 that she had written down before posting, which is the only reason nine was allowed to mean anything. Had she picked the number afterwards, nine enthusiastic comments would have felt like momentum.

The test cost her an evening of drawing and three weeks of patience. It saved $295 and a box of 100 pins. If you do decide to take money on a test like this rather than just names, the mechanics stop being a marketing question and become a legal one — the FTC's Mail, Internet, or Telephone Order Merchandise Rule (16 CFR Part 435 (opens in new tab)) governs what you must ship by when and what you owe if you cannot. The guide to taking preorders for handmade products covers how preorders, waitlists and backorders actually differ.

What the three tests had in common

None of them was expensive, and none of them produced a yes-or-no.

Every one of the three set the deciding number before the test ran. Nadia had three thresholds nine days out; Theo decided to count returning faces before market one; Ros published her number publicly, which made it very hard to move. This is the mechanism, and it works because it moves the judgment call to the only moment when you are neutral about the answer.

Every one measured something other than enthusiasm. Sell-through, minutes to sell out, asked-after-gone, returning buyers, committed names against a stated target. None of those is did people seem to like it, which is the metric a craft fair generates in abundance and which almost never changes anyone's decision.

And each test cost a small fraction of the commitment it was standing in for. Nadia risked 24 mugs against a season of production. Theo risked 40 bottles. Ros risked nothing but three weeks, against $295 and a drawer.

Two of the three verdicts were "yes, but differently" — a different price, a different role in the range. Only one was a no, and it was the cheapest one to receive.

Reading your own verdict

Decision tree titled What your test batch is telling you. From a root reading Test batch complete, units sold of units offered, three branches split on sell-through. Sold out, every unit gone, asks whether stock went before half the selling window or people asked after it ran out: Yes gives the verdict Underpriced, meaning same batch size, higher price, run it again and plan to have stock left at the end so you find the ceiling; No, it went late in the day, gives Passed on price, meaning repeat the batch at the same price and start counting how many buyers come back for a second one. Sold most of it, roughly 60 to 90 percent, asks how the repeat rate over two or more outings compares with your existing line: Close to it gives Product line, meaning scale to your next production constraint, the next kiln load, cook or cut of fabric, not to 100; Far below it gives Novelty, meaning keep the batches small and let it pull people to the table but do not build the next year of the range on it. Barely sold, under about a third, asks whether anyone stopped at the table at all: They walked past gives Wrong room, meaning venue or audience rather than the product, so retest somewhere else before changing anything; They picked it up then put it down gives You have your answer, a price or product objection, so stop here because this is the cheapest possible version of finding that out. A closing note reads: set the three thresholds, make more, stop, raise the price, before the batch leaves the house. A tree read after the fact is a rationalization.

Four results, four different next moves.

Sold out fast, and people asked afterwards. You were underpriced. Do not scale the batch — scale the price, keep the batch size, and deliberately plan to have stock left at the end so you can see where the ceiling is.

Sold most of it, at a pace, and some people came back. This is the boring result that actually means "make more." Scale to your next physical constraint, not to a hundred: the next kiln load up, the next cook size, the next fabric roll.

Sold most of it, and nobody came back. You have a novelty, and novelties are useful. Keep it in the range, keep the batches small, and let it do the job of stopping people at the table. Do not let it become the plan.

Barely sold. Work out whether anyone stopped at all. If people walked past without slowing, the problem is the venue or the audience and you should retest somewhere else before you touch the product. If people picked it up, turned it over, and put it down, you have your answer and it is a cheap one.

Whichever you get, run the winner against what you already sell before you commit the production time. In Ardent Seller that comparison lives in the sales reports — sell-through and margin per product, side by side — and it is a genuinely uncomfortable screen, because a new product that everybody complimented will sometimes sit below the boring thing you have made for four years. Better to find that out over 24 units than over a season.

The box in the garage was never about instinct. It was about deciding first and finding out afterwards. Make one kiln load, write three lines on a card, and let the card decide.

Ready to know what your test batch actually costs before you make it? Start free and build the batch as a recipe — materials, packaging and your own time — so the cost-per-unit on your scorecard is a figure rather than a guess.

Free resources

Free companion downloads if you want to run a test batch this month:


This article is provided for educational purposes only and does not constitute legal, regulatory, financial, or tax advice. Costs, pricing examples, and federal preorder and mail-order rules vary by jurisdiction and change frequently. Consult a qualified accountant, small-business advisor, or attorney before making financial or compliance decisions based on this content.

Frequently asked questions

Enough that the result is a number rather than a mood, and few enough that being wrong is survivable. Below about twelve units, a single enthusiastic buyer moves your sell-through by more than eight percentage points, so you are reading noise. Above about thirty, you have stopped testing and started producing. In practice the best answer is whatever your physical constraint already produces — one kiln load, one cook, one cut of fabric — because that is also the quantity you would have to repeat.

No. A discounted test tells you how people feel about the discount. One of the four things a test is for is finding out whether the price works, and you cannot learn that at a price you would never charge again. Test at the price you actually want, and treat slow sales as information rather than as a reason to mark down mid-test.

You move the test upstream. When tooling or a factory minimum forces the first commitment to be large — enamel pins, custom printed boxes, die-cut labels — you cannot test the product, so you test the demand signal instead: publish the finished design, name a number of committed buyers that would justify the order, and count. Set that number before you post it, because after you post it you will negotiate with yourself.

Partly. An interest test — a post, a listing, a waitlist — answers whether anyone wants it, which is the easiest of the four questions and the least informative. It cannot tell you what the thing costs to make at speed, and it cannot tell you whether people buy a second one. Use it when the first production commitment is genuinely expensive, and treat a pass as permission to run a real batch, not as a verdict.

At least two selling occasions. One outing answers whether anyone will buy it; only the second one starts to answer whether the same people come back, which is the question that separates a novelty from a product line. Three to five weeks is usually enough to get two or three markets in without the test quietly turning into your production schedule.

There is no universal number, which is exactly why you pick yours in advance. Write down three lines before the batch goes out: the figure at which you make more, the figure at which you stop, and the signal that says you were underpriced. Selling out in the first third of a market day is usually the third one, not the first — a fast sellout is a price signal, not a demand ceiling.