Hunch

Text classification in Excel with a formula

Updated · Edoardo Panichi, maker of Hunch

=HUNCH.PICK reads a cell of text and returns one option from the list you give it, spelled exactly as you wrote it. Add the same formula with a fourth argument to get a confidence number instead, which tells you which rows to check by hand.

The formula

=HUNCH.PICK(A2:A9, "sales lead: wants to buy or asks about pricing|support request: existing customer needs help|partnership: wants to collaborate or advertise|spam: irrelevant or promotional junk|other")

In Google Sheets: =HUNCHPICK(A2:A9, "sales lead: wants to buy or asks about pricing|support request: existing customer needs help|partnership: wants to collaborate or advertise|spam: irrelevant or promotional junk|other")

Options are separated by a pipe. Leave out the third argument and the default question is "Which option best describes this text?", which is right most of the time for a contact-form inbox, a feedback export, or any other column of short text.

=HUNCH.PICK(A2:A9, "sales lead: wants to buy or asks about pricing|support request: existing customer needs help|partnership: wants to collaborate or advertise|spam: irrelevant or promotional junk|other", "", "confidence")

In Google Sheets: =HUNCHPICK(A2:A9, "sales lead: wants to buy or asks about pricing|support request: existing customer needs help|partnership: wants to collaborate or advertise|spam: irrelevant or promotional junk|other", "", "confidence")

Same text, same options, empty third argument to keep the default question, and "confidence" as the fourth. The result is a number instead of a word: how sure the model is in the option it picked for column B, not a second opinion on the category itself.

Write categories the model does not have to guess at

Everything after a colon is a description the model reads; everything before it is the word that comes back in the cell. Use a description whenever two options could plausibly overlap, such as a sales question that is also a support question, or a press request that reads like spam on a first pass.

Always include other. Without it, a message about a job application or a broken link gets forced into the nearest category and quietly skews your counts. With it, the odd ones collect in one bucket you can read on a Friday.

Real outputs

Eight messages from a software company's contact form, with the model's answers as returned on 28 September 2026. Column C is the category, column D the confidence in that pick.

RowABC
1MessageCategoryConfidence
2Hi, do you have a plan for teams over 50 seats? Need pricing by Friday.sales lead1.00
3My export button just spins forever, using Chrome on Mac.support request1.00
4We run a newsletter with 40k subscribers in this space, want to talk sponsorship?partnership0.96
5CONGRATULATIONS you have been selected claim your prize nowspam1.00
6Love the product. Any chance of an API for our internal tool?support request0.52
7im tring to add my teammate but it says error code 402 pls helpsupport request1.00
8We write about productivity software, can we get a demo account for a review piece?partnership0.86
9heyother0.99

Row 5, an API question tacked onto a compliment, is the one to watch: the model picks support request but only at 0.52 confidence, exactly the ambiguous case a threshold is meant to catch, since the same line could as easily read as a sales question about integrations. Row 8, a single word ("hey"), is the opposite surprise: confidence is 0.99 even though the text carries almost no information, because none of the four named categories fit and other is an easy, confident call. Confidence measures how well the text matches the option chosen, not how much the model had to go on. The typo-heavy row 6 still resolves to support request at full confidence, which says the model is reading past the spelling, not matching keywords.

Find the rows worth a second look

=FILTER(A2:C9, C2:C9 < 0.6)

A FILTER on the confidence column pulls out exactly the rows a keyword rule would also have struggled with. Read those by hand; everything above the threshold is safe to trust and pivot on directly.

For a running count by category, either a PivotTable on column B or a plain COUNTIF works:

=COUNTIF(B2:B9, "spam")

Limits

Setup in three steps

  1. Sideload the add-in from hunchsheet.app/excel: Excel for Windows, Mac 2016 or later, and the web all support it.
  2. Save your key once: =HUNCH.SETKEY("hunch_..."), then delete the cell.
  3. Type the =HUNCH.PICK formula with your own options and fill it down the column.

See also

Questions

How many categories can one list have?
Up to 255. In practice, six to eight options plus other is easier for a human to review than twenty, even when the model could handle more.
Does the returned text match my option spelling exactly?
Yes, the label before the colon, character for character, including capitalization. That exactness is what lets a plain COUNTIF or SUMIF work against the column.
Do I need the confidence column every time?
No. Add it when a wrong pick is costly, such as routing a message to the wrong team, and skip it on a low-stakes tagging pass to save the extra credit per row.
What is a good confidence threshold to review?
Start at 0.6 and read what falls below it for a week. Confidence is specific to the option list; a five-option list and a two-option list will not sit at the same average.
Can one message belong to two categories?
HUNCH.PICK returns one. When a row can genuinely carry more than one tag at once, use HUNCH.MULTI instead, one yes/no question per tag.

Run it on your own column

Install the Excel add-in (two minutes, Windows, Mac and the web), save your key with =HUNCH.SETKEY, and point the formula at your data. 100 rows free to start, then $29 for 5,000 rows. Credits never expire and there is no subscription.

Reference: HUNCHPICK

More guides