← All articles

How we added an AI concierge to a static site

Infraheads logo

infraheads.com is a static site — hand-written HTML on S3, fronted by Cloudflare, with a two-minute prompt-to-deploy pipeline and no application server anywhere in the stack. So when we decided to add an AI concierge — a chat assistant that helps visitors pick the right Academy course or service and quietly captures a lead — the interesting question was how much server you actually need. The answer turned out to be "almost none."

The one rule you can't break

An LLM call needs a secret API key, and anything you ship to the browser is public — view-source, dev tools, a proxy, take your pick. Put the key in the widget's JavaScript and it gets scraped and billed against within days. That single fact rules out a genuinely "pure static" AI chatbot: you need somewhere server-side to hold the key. The whole design is about making that somewhere as small as possible.

The smallest backend that can keep a secret

That somewhere is one Cloudflare Worker. It runs on the same zone that already fronts the site, so a route hands it every request to /api/* — same origin, no CORS, no new domain:

routes = [
  { pattern = "infraheads.com/api/*", zone_name = "infraheads.com" }
]

The browser POSTs the conversation to /api/chat. The Worker prepends a system prompt grounded in our courses and services, calls Claude — the fast, inexpensive Haiku model — and streams the reply back as server-sent events. The API key lives only as an encrypted Worker secret: never in the repo, never in the S3 bucket, never in the browser. On Cloudflare's free plan the Worker itself costs nothing; the only real bill is Claude tokens, a fraction of a cent per conversation.

A widget that rides every page

The front end is a single self-contained file, chatbot.js — a floating launcher, a chat panel, and its own styles, all injected at runtime. It ships to every page the same way our analytics snippets do: a small build step stamps the loader into each HTML file, so there is no per-page markup to maintain. That is the same injection trick described in how we maintain infraheads.com.

Because the assistant is most useful at the moment of intent, we wired context-primed entry points next to the calls to action that already exist. A small "Ask the concierge" link sits beside each course's Enroll button and each consultation CTA; clicking it opens the panel already framed around that course or service, so the visitor never starts from a blank prompt.

The hidden goal: turn a chat into a lead

A concierge that only answers questions is a nice-to-have. The one that pays for itself captures a way to follow up. So the system prompt gives the assistant a quiet goal: once the conversation shows genuine interest, warmly ask for a name and email — once, never nagging — and the moment it has both, it emits a machine-readable directive as the last line of its reply:

<<lead:{"name":"...","email":"...","interest":"...","notes":"..."}>>

The widget strips that directive out of what the visitor sees and POSTs the captured fields to the exact mailbox our contact form already uses. No lead database, no CRM integration, no new plumbing — the concierge simply feeds the pipeline we had. The team gets an email; the visitor gets a "someone will follow up," plus a one-click path to enrollment or a discovery call.

Guardrails, because a public endpoint spends money

An unauthenticated endpoint that calls a paid API is a target. The Worker leans on cheap, layered defenses: it answers only requests from our own origin, applies a per-IP rate limit using Cloudflare's built-in rate-limiting binding (which stays on the free tier — no Durable Objects), caps message length and conversation history, and hard-limits the response size. Keep the blast radius small and the bill predictable.

What we reached for — and what we didn't

We deliberately skipped a lead store, a vector database, and a bespoke admin panel. The whole feature is a static widget plus one serverless function that borrows infrastructure we already ran. If abuse ever shows up we'll add a Turnstile challenge on the first message; if the concierge needs to answer deep questions about specific articles, we'll generate a small knowledge index from post metadata. Neither was needed to ship.

This is the pattern we like: solve the real problem with the least infrastructure that can hold a secret. It's the same instinct behind our self-improving access bot and the Lambda-less config API, and behind the platform work we do for clients — let the platform you already have do the job, and add only the smallest new piece the work actually requires.