No Intercom bill, no per-resolution fees — build your own AI support widget on Cloudflare's free tier.
If you've priced out an AI chatbot for your website recently, you already know the sticker shock. Intercom's Fin AI charges roughly a dollar per resolved conversation. Chatbase, Tidio, and most of the "no-code" options quietly cap your free plan at a few hundred messages a month before pushing you toward a $30–$100+ subscription. For a small blog, a SaaS side project, or a client site that doesn't get enterprise-level traffic, that math rarely makes sense.
In this guide, I'm going to show you how to build a free AI chatbot widget you can drop into any website with a single script tag — one that streams responses in real time, answers questions from your own FAQ using RAG (Retrieval Augmented Generation, meaning the AI looks up relevant facts before answering instead of guessing), and remembers a visitor's conversation across page reloads. We'll build it entirely on Cloudflare's free tier: Workers for the backend, Workers AI for the model, Vectorize for semantic search, and KV for session storage.
I've tested this exact stack on a low-traffic support page, and the free-tier limits (100,000 Worker requests/day, 10,000 AI "neurons"/day) were nowhere close to being an issue. If you're comfortable pasting commands into a terminal, you can have this live in under an hour — no credit card, no monthly invoice.
I'm Mostafa Amaan, and on Valley4Techs I write practical, hands-on guides for people who'd rather build the thing themselves than pay a subscription for it. Let's get into it.
What You'll Need Before You Start
This guide assumes you're comfortable with basic terminal commands and have a general idea of how APIs work, but you don't need prior Cloudflare experience. Here's the full checklist:
- A free Cloudflare account — sign up at dash.cloudflare.com. No credit card is required for the free tier.
- Node.js 18 or newer — installed on your machine. Download from nodejs.org.
You can verify with
node --version. - Basic JavaScript familiarity — enough to read and paste code. You won't be writing complex logic from scratch.
- A terminal or command prompt — on Windows, that's CMD or PowerShell; on macOS/Linux, any shell works.
- Your FAQ content ready — a short list of 3–10 question-and-answer pairs about your product or service.
- About an hour of focused time — the actual work is about 30 minutes of setup plus 15 minutes of testing, but having a full hour leaves room to debug if something unexpected comes up.
Why Build Your Own Chatbot Instead of Paying for Intercom or Chatbase?
I'm not against SaaS chatbot tools — if you're running a support team of ten agents, a platform like Intercom genuinely earns its price with ticketing, live handoff, and analytics baked in. But for the majority of website owners who just want visitors to get quick, accurate answers to common questions, you're paying for a lot of machinery you'll never touch. Here's how the free, self-hosted route actually compares:
| Option | Typical Cost | Setup Effort | You Own the Data |
|---|---|---|---|
| Intercom Fin AI | ~$0.99 per resolved chat + seat fees | Low (dashboard setup) | No |
| Chatbase / Tidio (free tier) | Free up to a low message cap, then $20–$100+/mo | Very low | Partially |
| Cloudflare Workers (this guide) | $0 for most low-to-mid traffic sites | Moderate (one-time build) | Yes, fully |
The trade-off is honest: you're spending an hour of setup time in exchange for full ownership and zero recurring cost. If you already understand the basics of how APIs work, you're more than qualified to follow along — you don't need prior experience with Cloudflare specifically.
What You're Building: The Architecture in Plain Terms
Before touching a terminal, it helps to see the moving parts. The whole project is two files talking to four Cloudflare services:
- A backend Worker — a small serverless function that receives chat messages, looks up relevant FAQ context, calls the AI model, and streams the reply back.
- A frontend widget script — a single embeddable JavaScript file that draws the chat bubble and window on any page it's loaded into.
- Workers AI — runs the actual language model (we'll use Meta's Llama 3) at the edge, so you're not paying OpenAI per token.
- Vectorize — a vector database. It stores your FAQ as numerical "embeddings" so the bot can find the most relevant answer by meaning, not just exact keyword matches. This is the RAG part.
- KV (Key-Value store) — Cloudflare's globally distributed storage, used here to remember each visitor's conversation between page loads.
None of these run on a server you manage. Cloudflare's edge network (300+ locations worldwide) runs your code close to whoever's visiting the site, which is also why response times stay snappy even on the free tier.
Step 1: Create the Project and Cloudflare Resources
You'll need a free Cloudflare account and Node.js 18 or newer installed. Start by scaffolding a Worker project:
Terminal
npm create cloudflare@latest support-widget cd support-widget npm install --save-dev tailwindcss wrangler
Pick javascript
when prompted, and choose "no" when asked to deploy immediately — we'll deploy once everything's ready. Next,
install Wrangler globally (Cloudflare's CLI) and log in:
Terminal
npm install -g wrangler wrangler login
Now create the two Cloudflare resources the project depends on — a Vectorize index for your FAQ, and a KV namespace for chat sessions:
Terminal
npx wrangler vectorize create support-faq --dimensions=768 --metric=cosine npx wrangler kv namespace create CHAT_HISTORY
Keep the ID printed by the second command handy — you'll paste it into the config file in the next step.
Step 2: Wire Everything Together in wrangler.jsonc
This config file tells your Worker which Cloudflare services it's allowed to use. Create
wrangler.jsonc
in your project root:
wrangler.jsonc
{
"name": "support-widget",
"main": "src/worker.js",
"compatibility_date": "2026-01-01",
"assets": { "directory": "./public", "binding": "ASSETS" },
"ai": { "binding": "AI" },
"vectorize": [
{ "binding": "FAQ_INDEX", "index_name": "support-faq" }
],
"kv_namespaces": [
{ "binding": "CHAT_HISTORY", "id": "PASTE_YOUR_KV_ID_HERE" }
]
}
ASSETS serves
your static widget files, AI gives the
Worker access to run models, FAQ_INDEX
connects to your vector database, and CHAT_HISTORY
is where visitor conversations get stored.
Step 3: Build the Backend Worker
This is where the actual logic lives — receiving a message, pulling relevant FAQ context, calling the model,
and streaming the answer back while saving it to KV. Create
src/worker.js:
src/worker.js
const SYSTEM_PROMPT = "You are a concise, friendly support assistant. Use the provided FAQ context when relevant. If you're unsure, say so instead of guessing.";
const SESSION_TTL = 60 * 60 * 24 * 14; // 14 days
// The widget almost always lives on a different domain than the site embedding it,
// so we reflect the real request origin instead of using a wildcard — a wildcard
// origin cannot be combined with credentialed (cookie-based) requests.
function corsHeaders(request) {
return {
"Access-Control-Allow-Origin": request.headers.get("Origin") || "*",
"Access-Control-Allow-Credentials": "true",
"Vary": "Origin",
"X-Content-Type-Options": "nosniff",
};
}
function getSessionId(request) {
const match = request.headers.get("Cookie")?.match(/widget_sid=([^;]+)/);
return match ? match[1] : null;
}
async function findRelevantFaq(env, question) {
const embedded = await env.AI.run("@cf/baai/bge-base-en-v1.5", { text: [question] });
if (!embedded?.data?.length) return "";
const matches = await env.FAQ_INDEX.query(embedded.data[0], { topK: 3, returnMetadata: true });
return matches.matches
.filter((m) => m.score > 0.6) // drop weak matches instead of feeding the model noise
.map((m) => `Q: ${m.metadata?.question}\nA: ${m.metadata?.answer}`)
.join("\n\n");
}
// Workers AI streams Server-Sent Events, but chunk boundaries don't respect line
// boundaries — a single "data: {...}" line can arrive split across two chunks.
// Buffering and only parsing complete lines avoids silently dropped tokens.
function createSseLineBuffer(onEvent) {
let buffer = "";
return {
push(text) {
buffer += text;
const lines = buffer.split("\n");
buffer = lines.pop();
for (const line of lines) {
if (line.startsWith("data: ") && !line.includes("[DONE]")) {
try { onEvent(JSON.parse(line.slice(6))); } catch (_) {}
}
}
},
};
}
async function handleChat(request, env) {
const cors = corsHeaders(request);
const { message } = await request.json();
if (!message?.trim()) {
return new Response(JSON.stringify({ error: "Message is required" }), { status: 400, headers: { "Content-Type": "application/json", ...cors } });
}
let sessionId = getSessionId(request);
let session = sessionId ? await env.CHAT_HISTORY.get(sessionId, "json") : null;
// needsCookie must track whether we actually generated a NEW session id —
// not just whether a cookie was present. If the cookie survived but the KV
// entry expired, the old code below would silently start a fresh session
// without ever re-issuing a cookie, breaking memory permanently.
let needsCookie = false;
if (!session) {
sessionId = crypto.randomUUID();
session = { history: [] };
needsCookie = true;
}
session.history.push({ role: "user", content: message.trim() });
const faqContext = await findRelevantFaq(env, message);
const promptMessages = [
{ role: "system", content: SYSTEM_PROMPT + (faqContext ? `\n\nRelevant FAQ:\n${faqContext}` : "") },
...session.history.slice(-8),
];
const modelStream = await env.AI.run("@cf/meta/llama-3-8b-instruct", { messages: promptMessages, stream: true });
let fullReply = "";
const decoder = new TextDecoder();
const lineBuffer = createSseLineBuffer((chunk) => { fullReply += chunk.response || ""; });
const { readable, writable } = new TransformStream({
transform(chunk, controller) {
lineBuffer.push(decoder.decode(chunk, { stream: true }));
controller.enqueue(chunk);
},
async flush() {
if (fullReply) {
session.history.push({ role: "assistant", content: fullReply });
await env.CHAT_HISTORY.put(sessionId, JSON.stringify(session), { expirationTtl: SESSION_TTL });
}
},
});
modelStream.pipeTo(writable).catch((err) => console.error("Stream error:", err));
const headers = { "Content-Type": "text/event-stream", ...cors };
// SameSite=None + Secure is required (not Lax) because this is a cross-site
// request from the host page's origin to the Worker's origin. Lax cookies
// are withheld on cross-site fetch calls, which silently breaks memory.
if (needsCookie) headers["Set-Cookie"] = `widget_sid=${sessionId}; Path=/; HttpOnly; SameSite=None; Secure; Max-Age=${SESSION_TTL}`;
return new Response(readable, { headers });
}
async function handleHistory(request, env) {
const cors = corsHeaders(request);
const sessionId = getSessionId(request);
const session = sessionId ? await env.CHAT_HISTORY.get(sessionId, "json") : null;
return new Response(JSON.stringify({ history: session?.history || [] }), { headers: { "Content-Type": "application/json", ...cors } });
}
async function handleSeed(request, env) {
const cors = corsHeaders(request);
const providedKey = request.headers.get("X-Seed-Key");
if (!env.SEED_KEY || providedKey !== env.SEED_KEY) {
return new Response(JSON.stringify({ error: "Unauthorized" }), { status: 401, headers: { "Content-Type": "application/json", ...cors } });
}
const { faq } = await request.json(); // faq: [{ question, answer }, ...]
if (!Array.isArray(faq) || !faq.length) {
return new Response(JSON.stringify({ error: "Provide a non-empty faq array" }), { status: 400, headers: { "Content-Type": "application/json", ...cors } });
}
const vectors = [];
for (const entry of faq) {
const embedded = await env.AI.run("@cf/baai/bge-base-en-v1.5", { text: [entry.question] });
vectors.push({
id: crypto.randomUUID(),
values: embedded.data[0],
metadata: { question: entry.question, answer: entry.answer },
});
}
await env.FAQ_INDEX.upsert(vectors);
return new Response(JSON.stringify({ seeded: vectors.length }), { headers: { "Content-Type": "application/json", ...cors } });
}
export default {
async fetch(request, env) {
const path = new URL(request.url).pathname;
const cors = corsHeaders(request);
if (request.method === "OPTIONS") {
return new Response(null, { headers: { ...cors, "Access-Control-Allow-Methods": "GET,POST,OPTIONS", "Access-Control-Allow-Headers": "Content-Type, X-Seed-Key" } });
}
if (path === "/api/chat" && request.method === "POST") return handleChat(request, env);
if (path === "/api/history") return handleHistory(request, env);
if (path === "/api/seed" && request.method === "POST") return handleSeed(request, env);
return env.ASSETS.fetch(request);
},
};
Access-Control-Allow-Origin: *
combined with a credentialed (cookie-carrying) request is rejected by every modern browser — and since your
widget's domain is almost never the same as the site embedding it, this is a cross-site request by default.
On top of that, a SameSite=Lax
cookie is withheld on cross-site fetch calls entirely. Get either of these wrong and the chatbot will appear
to work in your first message, then quietly forget everything on the next one — no error, no console warning,
just a bot with amnesia. That's why the code above reflects the real request origin and uses
SameSite=None; Secure.
A few more details worth calling out. The session cookie is HttpOnly,
which stops client-side scripts from reading it — a small but important security habit for anything that
tracks a visitor across requests. The FAQ lookup filters out weak matches below a similarity score of 0.6, so
a completely unrelated question doesn't get padded with irrelevant "context" that confuses the model. And the
SSE parser buffers partial lines across chunk boundaries instead of parsing each chunk in isolation — without
that buffer, streamed responses can drop random words whenever a JSON line happens to split across two network
packets, which is the kind of bug that only shows up intermittently and is miserable to debug after the fact.
Finally, notice that needsCookie
is set whenever a new session is created — not just when the incoming request had no cookie at all.
If a visitor's cookie survives but their KV entry has expired (after the 14-day TTL), the original approach
would silently start a fresh session without ever re-issuing a cookie, leaving that visitor's history
unrecoverable on every future visit. It's the kind of edge case that passes a quick manual test and only
breaks two weeks into production.
Step 4: Set Up Tailwind CSS for the Widget
Since this widget gets embedded on someone else's site, using a full CSS framework with scoped utility classes
avoids clashing with the host page's existing styles. Create
tailwind.config.js:
tailwind.config.js
module.exports = {
content: ["./public/**/*.{html,js}"],
darkMode: "class",
theme: { extend: {} },
};
Then add a source stylesheet at
src/styles.css
with the three standard Tailwind directives:
src/styles.css
@tailwind base; @tailwind components; @tailwind utilities;
Then add these scripts to
package.json:
package.json (scripts section)
"scripts": {
"build:css": "npx tailwindcss -i ./src/styles.css -o ./public/widget.css --minify",
"dev": "npm run build:css && wrangler dev",
"deploy": "npm run build:css && wrangler deploy"
}
Step 5: Build the Embeddable Widget Script
This is the file that any website will load with a single <script>
tag. It builds the chat bubble, handles opening and closing the window, and streams the AI's response into the
page. Create public/widget.js:
public/widget.js
(function () {
const config = {
baseUrl: window.SUPPORT_WIDGET_URL || "",
title: window.SUPPORT_WIDGET_TITLE || "Support",
greeting: window.SUPPORT_WIDGET_GREETING || "Hi! What can I help you with?",
};
let messages = [];
let isOpen = false;
let isSending = false;
// Message content comes from the visitor AND from the model — neither is
// trusted. Without escaping, typing "<img src=x onerror=alert(1)>" into the
// chat box would execute arbitrary script on the host page.
function escapeHtml(str) {
return str
.replace(/&/g, "&")
.replace(/innerHTML.
If you insert that text raw, a visitor typing something like an
<img
onerror=...> tag into the chat box gets executed as real HTML on whatever site the widget is
embedded on — a stored cross-site scripting vulnerability. The
escapeHtml()
helper above is not optional polish; skip it and you've shipped a security hole on every site that installs
your widget.
The whole thing is wrapped in an IIFE (a function that runs immediately and doesn't leak variables into the
host page's global scope) — an important habit for any script meant to be dropped onto someone else's site.
The streaming loop in sendMessage()
updates the same message bubble in place as tokens arrive, rather than re-rendering the whole chat window on
every chunk — that's what produces the smooth "typing" effect instead of visible flicker. The
try/catch
around the fetch call means a dropped connection shows the visitor a friendly fallback message instead of a
chat window that just hangs forever with no feedback — something the bare-bones version of this pattern
almost always skips.
Step 6: Test Locally, Then Deploy
One file is still missing: since env.ASSETS.fetch()
serves whatever's in ./public,
visiting the root URL right now would just 404 — there's no page there yet. Add a minimal demo page so
wrangler dev
actually has something to show you. Create
public/index.html:
public/index.html
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<title>Support Widget — Local Demo</title>
</head>
<body style="font-family: sans-serif; padding: 40px; max-width: 640px; margin: 0 auto;">
<h1>Local demo page</h1>
<p>This page only exists so the dev server has something to render at the root.
The chat bubble should appear in the bottom-right corner — click it to test.</p>
<script>
window.SUPPORT_WIDGET_URL = ""; // same-origin while testing locally
window.SUPPORT_WIDGET_TITLE = "Ask Us Anything";
</script>
<script src="/widget.js"></script>
</body>
</html>
Run the dev server and open it in your browser:
Terminal
npm run dev # opens at http://localhost:8787
Click the chat bubble in the corner and send a test message — you should see it stream back token by token. If nothing happens, open your browser's console first; a CORS or cookie error will show up there immediately.
SameSite=None; Secure,
which is required for cross-site requests in production. However, modern browsers reject the
Secure
flag on plain http://localhost
connections. This means the session cookie will not be set during local testing with
wrangler dev,
and the chatbot will appear to forget the conversation after each message. This is expected — the cookie works
correctly once deployed to Cloudflare's HTTPS endpoint. To test locally with full session persistence, use
wrangler dev --remote
(which uses your live Cloudflare account with HTTPS) or tunnel through a tool like
ngrok.
Once you're happy with it locally, deploy with a single command:
Terminal
npm run deploy
You'll get a live URL that looks like
support-widget.YOUR-SUBDOMAIN.workers.dev.
The bot won't be able to answer FAQ questions yet, though — the Vectorize index is still empty. That's what
the next step fixes.
Seed Your FAQ Data into Vectorize
Before the chatbot can answer real questions, you need to convert your FAQ into embeddings and store them in
the vector index. The handleSeed()
function already added to worker.js
(in the code block from Step 3, under the
/api/seed
route) handles this — it accepts a list of question/answer pairs, embeds each question, and upserts the
result into Vectorize. It's protected by a secret key so a stranger can't overwrite your FAQ.
Set that secret before deploying:
Terminal
npx wrangler secret put SEED_KEY # paste any random string when prompted — you'll reuse it below
Then call the endpoint once with your FAQ content:
Terminal
curl -X POST https://support-widget.YOUR-SUBDOMAIN.workers.dev/api/seed \
-H "Content-Type: application/json" \
-H "X-Seed-Key: YOUR_SECRET_HERE" \
-d '{
"faq": [
{ "question": "What are your shipping times?", "answer": "Orders ship within 2 business days and arrive in 3-7 days depending on location." },
{ "question": "Do you offer refunds?", "answer": "Yes, full refunds are available within 30 days of purchase." },
{ "question": "How do I contact support?", "answer": "Email support@example.com or use this chat widget." }
]
}'
A successful response returns
{"seeded": 3}
— confirming three FAQ entries were embedded and stored. Re-run this same call any time your FAQ content
changes; there's no need to clear the index first since each entry gets a fresh unique ID.
Step 7: Embed the Widget on Any Website
This is the whole point — adding your chatbot to any site, including one built on a completely different
stack, takes two lines before the closing
</body>
tag:
HTML
<script> window.SUPPORT_WIDGET_URL = "https://support-widget.YOUR-SUBDOMAIN.workers.dev"; window.SUPPORT_WIDGET_TITLE = "Ask Us Anything"; </script> <script src="https://support-widget.YOUR-SUBDOMAIN.workers.dev/widget.js"></script>
That's it. Because the widget is a self-contained script with its own styles, it won't fight with your site's existing CSS — it'll work the same whether the host page is built with WordPress, a static site generator, or a custom Shopify theme.
5 Mistakes I See People Make With Self-Hosted Chatbots
- Skipping the FAQ context entirely. Without RAG, you're just running a generic chat model that knows nothing about your product. It'll hallucinate confidently. Even a handful of well-written Q&A pairs makes a noticeable difference in accuracy.
- Not capping conversation history. If you send the entire chat history to the model on every message, both your latency and your neuron usage climb fast. Trim to the last 6–10 messages, which is plenty of context for support-style conversations.
-
Using a wildcard CORS origin with cookies. Your widget is hosted on a different domain
than the page it's embedded in, and a wildcard
Access-Control-Allow-Origin: *simply doesn't work once cookies are involved — browsers block it outright. We covered the fix (reflecting the real origin, plusSameSite=None) back in Step 3, but it's worth repeating because it's the single most common reason a self-hosted chatbot "loses" the conversation after the first message. - No fallback for "I don't know." Bake an instruction into your system prompt that tells the model to admit uncertainty rather than invent an answer. A wrong answer damages trust far more than an honest "let me connect you with a human."
- Never re-seeding Vectorize after FAQ updates. Your live chatbot only knows what's in the vector index — if you update your FAQ page but forget to re-embed the new content, the bot keeps answering from stale data indefinitely.
Cloudflare's Free Tier: What You Actually Get for $0
Before you commit to this route, it's worth knowing exactly where the free tier ends. From my own testing, these limits are generous enough for the vast majority of small-to-medium sites:
- Workers: 100,000 requests per day
- Workers AI: 10,000 "neurons" per day (roughly a few hundred short conversations, depending on message length)
- Vectorize: 5 million vector query/insert operations per month
- KV: 100,000 reads and 1,000 writes per day
If you outgrow these — congratulations, your site has real traffic — Cloudflare's paid tiers scale up without requiring you to rebuild anything, since you're already on their infrastructure.
Where to Take This Project Next
Once the basic widget is live, a few upgrades are worth exploring. You could swap Llama 3 for a different open model to compare response quality — our Qwen vs GPT vs Gemini comparison is a good starting point for understanding the trade-offs between models. If you'd rather orchestrate the chatbot's logic visually instead of hand-coding it, our n8n automation guide covers a no-code alternative worth comparing against this hand-built approach. And if the difference between an "AI agent" and a simple chatbot has you curious, we cover that distinction in our guide on building a free team of AI agents for your website.
Want to version-control your Worker code and collaborate with others on it? Our Git and GitHub beginner's guide covers exactly that, and it pairs well with pushing this project to a repository before you start customizing it further.
Troubleshooting Common Issues
Even with correct code, things can go wrong. Here are the most common issues and how to fix them:
Customizing the Widget's Look and Feel
The widget uses Tailwind utility classes, which makes customization straightforward. Here are a few common changes:
- Change the accent color — Replace every
bg-emerald-600inwidget.jswith your brand color, e.g.bg-blue-600. The same applies totext-emerald-600and the green gradient in the border. - Move the chat bubble — Adjust the
bottom-6 right-6classes. For a left-aligned widget, usebottom-6 left-6instead. - Change the widget size — Modify
w-96(width) andh-[560px](height) in the panel div. - Replace the bubble icon — Change the
💬emoji in the toggle button to your own icon or SVG. - Add your own CSS overrides — Add styles to
src/styles.cssbefore the Tailwind directives, or add a second stylesheet afterwidget.css.
Tracking Chatbot Usage and Performance
Once your widget is live, you'll want to know how it's being used. Cloudflare's dashboard gives you basic metrics (request counts, error rates), but for deeper insight, consider these approaches:
- Cloudflare Analytics dashboard — Visit the Cloudflare Dashboard → Workers & Pages → your Worker. You'll see request volume, CPU time, and error rates out of the box — no extra code needed.
- Log into KV events — Add a small logging step inside
handleChat()that writes anonymized usage data (question asked, whether a FAQ match was found, response time) to a separate KV namespace. This lets you analyze which questions are most common and whether the bot is answering them well. - Manual review via
wrangler tail— Runnpx wrangler tailto watch live logs from your deployed Worker. Useful for debugging real-time issues. - User feedback integration — Add a "Was this helpful?" thumbs-up/thumbs-down button after each bot response. Store the feedback in KV to measure satisfaction over time.
Going to Production: What's Missing
The widget as built above works, but a production deployment should add a few more layers:
- Custom domain — Instead of the
*.workers.devsubdomain, point your own domain (e.g.widget.yourdomain.com) to the Worker. Add arouteskey inwrangler.jsonc:"routes": [{ "pattern": "widget.yourdomain.com", "custom_domain": true }]. - Rate limiting — Add basic rate limiting to prevent abuse. Cloudflare's Rate Limiting example shows how to add a counter in KV and reject requests after a threshold.
- Error monitoring — Integrate with Workers Observability to get alerts when error rates spike.
- Human handoff — For questions the bot can't answer, add a "Talk to a human" button that sends the conversation transcript to your support email or a Slack webhook.
Privacy, GDPR, and Compliance Considerations
If your website serves visitors in the EU, UK, or California, a chatbot that stores conversation history needs a few compliance basics:
- Cookie consent — The
widget_sidcookie is a strictly necessary cookie (it powers the conversation feature), but your cookie consent banner should still disclose it. Update your privacy policy to mention the cookie and its 14-day retention period. - Data retention — The 14-day TTL on KV entries means conversations are automatically deleted. If you need longer retention for analytics, anonymize the data and store it separately with a clear retention policy.
- Right to deletion — Add a
DELETE /api/historyendpoint that clears the visitor's KV entry, allowing users to exercise their right to be forgotten under GDPR Article 17. - Data processing addendum — If you're using this widget for a client site, Cloudflare acts as a data processor. Your clients may need a DPA (Data Processing Addendum) with Cloudflare, which is available in your Cloudflare dashboard under Account → Agreements.
What Happens When You Outgrow the Free Tier
Cloudflare's free tier is generous, but it's worth planning for growth. Here's what the paid tiers look like and when you'd need them:
| Service | Free Limit | Paid Starts At | When You'd Need It |
|---|---|---|---|
| Workers | 100k req/day | ~$5/month (Workers Paid) | ~3,000 visitors/day |
| Workers AI | 10k neurons/day | ~$0.001/1k neurons | ~500 conversations/day |
| Vectorize | 5M operations/mo | Usage-based | ~50k FAQ queries/day |
| KV | 100k reads + 1k writes/day | ~$0.50/month for 1M reads | ~3k visitors/day |
Even the Workers Paid plan at $5/month is dramatically cheaper than any SaaS chatbot for similar traffic levels. The key advantage is that pricing is usage-based — you only pay for what you actually use, and your existing code works without modification.
Final Thoughts
What you've built here isn't a toy — it's a legitimate, production-capable alternative to tools that routinely charge hundreds of dollars a month. The pieces (a serverless function, a vector database, and a key-value store) are the same building blocks powering much larger AI products; you're just running them at a scale that happens to fit inside Cloudflare's free tier.
The parts most tutorials skip — trimming conversation history, handling CORS properly, giving the model permission to say "I don't know" — are exactly the details that separate a demo from something you'd actually trust on a live website. Get those right, and the difference between this and a $99/month SaaS widget mostly comes down to who owns the setup script.
If you build this out, I'd genuinely like to know how the free-tier limits hold up on your traffic — drop a comment once it's live.
Liked the hands-on approach?
Join hundreds of subscribers and get practical AI, security, and networking guides — projects, not theory — delivered to your inbox.
Yes, Subscribe Me! ✉️🔒 No spam, ever. We respect your inbox.
We'd love to hear your thoughts! Leave a comment below
and share your experience or questions.