Six thousand products, described once and found in search
A headless storefront rebuild with attribute extraction and drafted product copy, where a merchandiser approves everything before it publishes.
- Products with a full attribute set
- 6,000
- Descriptions approved by a person
- 100%
- Behind filters, search and markup
- 1 source
The client
A direct-to-consumer home and lifestyle retailer
Named clients are withheld under the confidentiality agreements we work to. Sector, scale and system detail are published with permission.
- Sector
- Online retail
- Scale
- About 6,000 live products, two merchandisers
- Duration
- 16 weeks, with catalogue migration running alongside the build
- Team
- Frontend engineer, Backend engineer, AI engineer, SEO specialist, Merchandising lead, client side, Project manager
What was wrong.
Half the catalogue had no description beyond a supplier code, and the other half repeated the same forty words. Product pages were slow on a mobile connection, filters behaved differently in each category, and the site's own search missed products that were plainly in stock.
Two merchandisers could not write 3,000 descriptions, and nobody wanted a catalogue of unread machine text going live. The requirement was a system that does the attribute work and the drafting, and still puts a person in front of anything that publishes.
The approach, step by step.
- 01
Fix the data model first
We built one attribute schema per category: material, dimensions, care, room, finish. Everything after this step depends on a product having attributes rather than a paragraph.
- 02
Extract before you generate
Attributes are pulled from supplier sheets, legacy copy and image labels into that schema. Where a value is missing the field stays empty, because an invented dimension is worse than a blank one.
- 03
Draft, then approve
A model drafts the description from the attributes alone, in the brand's tone and inside a length budget set per category. Merchandisers see the draft beside its source attributes and publish, edit or reject.
- 04
Rebuild the storefront around it
A headless front end renders from the same attribute data, so filters, product schema markup and the on-site search index are generated rather than maintained by hand.
- 05
Measure what search can see
Structured data, canonical rules and a generated sitemap were checked against index coverage every week through the first quarter, and category templates were corrected where pages were being crawled and skipped.
The architecture we shipped.
Every layer below exists in the running system. Nothing here is a reference diagram.
- Product data
- One versioned attribute schema per category, with required and optional fields declared up front.
- Extraction pipeline
- Supplier sheets, legacy copy and image labels mapped into attributes, with a confidence per field and no silent defaults.
- Drafting service
- A prompt per category writes from attributes only, inside a length budget and against a banned-claims list for regulated wording.
- Approval queue
- Draft, source attributes and diff in one screen. Nothing publishes without a merchandiser acting on it.
- Commerce back end
- Headless catalogue and orders, with published content served through a cache invalidated per product.
- Storefront
- Server-rendered category and product pages with generated product markup, image sizing and per-route caching.
- Search and reporting
- The on-site search index is rebuilt from attributes, alongside a weekly crawl and index-coverage report.
What it does now.
Read from the system itself. No revenue claims, no multiples.
Products with a full attribute set
One schema per category, filled from source data rather than free text.
Descriptions approved by a person
Drafts wait in a queue; nothing publishes unread.
Behind filters, search and markup
The same attributes drive all three, so they cannot drift apart.
Weekly, on index coverage
Crawled, indexed and skipped pages, checked against the generated sitemap.
What it runs on.
- Next.js
- TypeScript
- Headless commerce API
- PostgreSQL
- Per-category LLM drafting
- Typesense
- Cloudflare CDN
- Playwright regression tests
- Schema.org structured data
The drafts do not publish themselves, and that is the part I like. I read them at ten times the speed I could write them.
Attributed by role and sector only, at the client's request.
What happens next.
Regional language descriptions drafted from the same attributes and approved through the same queue.
Services behind this build
Sector
Retail & MallsAbout 6,000 live products, two merchandisers. The pattern transfers; the domain detail is rebuilt for every client.
Have a system that should work like this one?
We will walk your process, tell you what is worth automating, and scope the first version that can be measured.