A Checklist for Choosing a Site Search Vendor for a Large Catalog
Score vendors on your catalog, not their demo
Every site search demo looks sharp because it runs on a curated dataset and friendly queries. Your catalogue is 100,000-plus SKUs from suppliers who each name things their own way, and your shoppers type "warm waterproof boots" and full sentences. This site search vendor checklist is built around whether a platform actually handles that, and whether it keeps up as the catalogue grows.
1. Relevance, evaluated on your data
- Will they index a sample of your catalogue and run your query logs, not their demo set?
- Zero-results rate before and after, and performance on the hard queries: natural language, attribute phrases ("boots for wide feet under 100"), misspellings, brand-versus-generic, supplier synonyms.
- Can you A/B a slice of live traffic before committing?
2. Semantic, lexical, or hybrid
- Is it true semantic retrieval (embeddings), a synonym-expansion layer bolted onto keyword search, or a genuine hybrid?
- Does it keep exact matching for SKUs, part numbers, and model names — the queries where "close" is wrong?
3. Index freshness
- How fast does a new or changed SKU become searchable — seconds, or a nightly batch?
- Does search reflect live price and stock, or a snapshot? Ranking a sold-out item first erodes trust in the whole feature.
- A supplier feed adding thousands of SKUs a week — can the indexing pipeline keep pace without falling behind?
4. Merchandising control
- Boost, bury, pin, filter, and redirect — can a non-technical merchandiser do these without an engineering ticket?
- Where do relevance rules live, who owns them, and is there an audit log of changes?
5. Generative features
- Ranked results only, or grounded natural-language answers and recommendations?
- Is that a higher pricing tier, and does it call an external model you're then also paying for?
6. Latency at your volume
- p95 response time under type-ahead thresholds, measured at your query rate and catalogue size.
- How latency behaves as the catalogue grows toward your two-year projection.
7. Pricing and lock-in
- Per request, per record, per seat — and how the bill behaves at 2x and 5x your current catalogue and query volume.
- Is the AI-relevance tier a multiple of the base search price?
- Can you export the index, the relevance rules, the synonym lists, and the analytics if you leave?
8. Data handling
- Where catalogue and query data go, which sub-processors touch them, and the retention terms.
How to weight it
A smaller, stable catalogue makes sections 1 and 4 the decision — relevance and control. A large catalogue on a fast supplier feed makes sections 3, 6, and 7 dominant — freshness and pricing at scale are where these deployments actually strain. Our semantic search project weighed a managed platform against a self-hosted engine and a custom stack precisely on those axes: a 120,000-SKU catalogue growing by thousands a week, where per-request-plus-per-record pricing and index freshness mattered more than out-of-the-box relevance.
Red flags
- The evaluation is on their data. If they won't run your catalogue and query logs, they're hiding the relevance and the zero-results rate.
- "Relevance tunes itself." Some of it does; your promo priorities and discontinued lines don't. Ask exactly where business overrides live.
- Vague on freshness with a fast feed. "How long until a new SKU is searchable" should get a number.
- Pricing not quoted at your real record count and at 3x volume.
- An index and rules you can't export. That's lock-in by design.
Where this checklist misleads
- Feature-counting. A vendor can tick every box with capabilities you'll never switch on. Weight by your catalogue and query mix.
- Demo relevance. A curated demo tells you little; insist on your own catalogue and a live A/B.
- The smallest catalogues. A few thousand stable SKUs may be well served by keyword search plus a synonym list, with no vendor at all.
FAQ
Build or buy — which does this point to? Either. Score vendors with it, or use the same questions as requirements if you build.
What's the single most important item? Relevance measured on your own catalogue and query logs. A demo on their data is not evidence.
How long should evaluation take? Long enough to index your catalogue, replay real query logs, and run a live A/B — weeks, not a one-hour demo.
ISTRALLEN helps retail teams evaluate search platforms and builds custom stacks when platform pricing or control doesn't fit; see AI for Retail.