# Text classification in Excel with a formula

`=HUNCH.PICK` reads a cell of text and returns one option from the list you give it, spelled exactly as you wrote it. Add the same formula with a fourth argument to get a confidence number instead, which tells you which rows to check by hand.

Updated 28 September 2026. Source: https://hunchsheet.app/formulas/classify-text-excel

## The formula

```
=HUNCH.PICK(A2:A9, "sales lead: wants to buy or asks about pricing|support request: existing customer needs help|partnership: wants to collaborate or advertise|spam: irrelevant or promotional junk|other")
```
In Google Sheets: `=HUNCHPICK(A2:A9, "sales lead: wants to buy or asks about pricing|support request: existing customer needs help|partnership: wants to collaborate or advertise|spam: irrelevant or promotional junk|other")`

Options are separated by a pipe. Leave out the third argument and the default question is "Which option best describes this text?", which is right most of the time for a contact-form inbox, a feedback export, or any other column of short text.

```
=HUNCH.PICK(A2:A9, "sales lead: wants to buy or asks about pricing|support request: existing customer needs help|partnership: wants to collaborate or advertise|spam: irrelevant or promotional junk|other", "", "confidence")
```
In Google Sheets: `=HUNCHPICK(A2:A9, "sales lead: wants to buy or asks about pricing|support request: existing customer needs help|partnership: wants to collaborate or advertise|spam: irrelevant or promotional junk|other", "", "confidence")`

Same text, same options, empty third argument to keep the default question, and `"confidence"` as the fourth. The result is a number instead of a word: how sure the model is in the option it picked for column B, not a second opinion on the category itself.

## Write categories the model does not have to guess at

Everything after a colon is a description the model reads; everything before it is the word that comes back in the cell. Use a description whenever two options could plausibly overlap, such as a sales question that is also a support question, or a press request that reads like spam on a first pass.

> Tip: Always include other. Without it, a message about a job application or a broken link gets forced into the nearest category and quietly skews your counts. With it, the odd ones collect in one bucket you can read on a Friday.

## Real outputs

Eight messages from a software company's contact form, with the model's answers as returned on 28 September 2026. Column C is the category, column D the confidence in that pick.

| Message | Category | Confidence |
| --- | --- | --- |
| Hi, do you have a plan for teams over 50 seats? Need pricing by Friday. | sales lead | 1.00 |
| My export button just spins forever, using Chrome on Mac. | support request | 1.00 |
| We run a newsletter with 40k subscribers in this space, want to talk sponsorship? | partnership | 0.96 |
| CONGRATULATIONS you have been selected claim your prize now | spam | 1.00 |
| Love the product. Any chance of an API for our internal tool? | support request | 0.52 |
| im tring to add my teammate but it says error code 402 pls help | support request | 1.00 |
| We write about productivity software, can we get a demo account for a review piece? | partnership | 0.86 |
| hey | other | 0.99 |

Row 5, an API question tacked onto a compliment, is the one to watch: the model picks support request but only at 0.52 confidence, exactly the ambiguous case a threshold is meant to catch, since the same line could as easily read as a sales question about integrations. Row 8, a single word ("hey"), is the opposite surprise: confidence is 0.99 even though the text carries almost no information, because none of the four named categories fit and other is an easy, confident call. Confidence measures how well the text matches the option chosen, not how much the model had to go on. The typo-heavy row 6 still resolves to support request at full confidence, which says the model is reading past the spelling, not matching keywords.

## Find the rows worth a second look

```
=FILTER(A2:C9, C2:C9 < 0.6)
```

A FILTER on the confidence column pulls out exactly the rows a keyword rule would also have struggled with. Read those by hand; everything above the threshold is safe to trust and pivot on directly.

For a running count by category, either a PivotTable on column B or a plain COUNTIF works:

```
=COUNTIF(B2:B9, "spam")
```

## Limits

- A list can hold up to 255 options. Give the full list rather than a shortlist; long option lists cost a few more tokens per row but no more credits.
- The option text comes back character for character, which is what makes COUNTIF, FILTER and pivots line up against it.
- Confidence is its own formula call, so a category-plus-confidence column is two credits per row, not one.
- Cells calculating together travel in requests of up to 40 rows, and results are cached for six hours inside Excel.

## Setup in three steps

1. Sideload the add-in from [hunchsheet.app/excel](https://hunchsheet.app/excel): Excel for Windows, Mac 2016 or later, and the web all support it.
2. Save your key once: `=HUNCH.SETKEY("hunch_...")`, then delete the cell.
3. Type the `=HUNCH.PICK` formula with your own options and fill it down the column.

## See also

- [HUNCH.PICK reference](https://hunchsheet.app/docs/hunchpick) for the full argument list, including `"probabilities"`.
- [The same pattern in Google Sheets](https://hunchsheet.app/formulas/classify-text-google-sheets-formula).
- [Categorize support tickets in Excel](https://hunchsheet.app/formulas/categorize-support-tickets-excel), a worked example with its own category column.

## Questions

### How many categories can one list have?

Up to 255. In practice, six to eight options plus other is easier for a human to review than twenty, even when the model could handle more.

### Does the returned text match my option spelling exactly?

Yes, the label before the colon, character for character, including capitalization. That exactness is what lets a plain COUNTIF or SUMIF work against the column.

### Do I need the confidence column every time?

No. Add it when a wrong pick is costly, such as routing a message to the wrong team, and skip it on a low-stakes tagging pass to save the extra credit per row.

### What is a good confidence threshold to review?

Start at 0.6 and read what falls below it for a week. Confidence is specific to the option list; a five-option list and a two-option list will not sit at the same average.

### Can one message belong to two categories?

HUNCH.PICK returns one. When a row can genuinely carry more than one tag at once, use [HUNCH.MULTI](https://hunchsheet.app/docs/hunchmulti) instead, one yes/no question per tag.

