📁 last tech Posts

Build a Free AI Chatbot for Your Website with Cloudflare

Build a free AI chatbot widget for your website using Cloudflare Workers

No Intercom bill, no per-resolution fees — build your own AI support widget on Cloudflare's free tier.

If you've priced out an AI chatbot for your website recently, you already know the sticker shock. Intercom's Fin AI charges roughly a dollar per resolved conversation. Chatbase, Tidio, and most of the "no-code" options quietly cap your free plan at a few hundred messages a month before pushing you toward a $30–$100+ subscription. For a small blog, a SaaS side project, or a client site that doesn't get enterprise-level traffic, that math rarely makes sense.

In this guide, I'm going to show you how to build a free AI chatbot widget you can drop into any website with a single script tag — one that streams responses in real time, answers questions from your own FAQ using RAG (Retrieval Augmented Generation, meaning the AI looks up relevant facts before answering instead of guessing), and remembers a visitor's conversation across page reloads. We'll build it entirely on Cloudflare's free tier: Workers for the backend, Workers AI for the model, Vectorize for semantic search, and KV for session storage.

I've tested this exact stack on a low-traffic support page, and the free-tier limits (100,000 Worker requests/day, 10,000 AI "neurons"/day) were nowhere close to being an issue. If you're comfortable pasting commands into a terminal, you can have this live in under an hour — no credit card, no monthly invoice.

I'm Mostafa Amaan, and on Valley4Techs I write practical, hands-on guides for people who'd rather build the thing themselves than pay a subscription for it. Let's get into it.

What You'll Need Before You Start

This guide assumes you're comfortable with basic terminal commands and have a general idea of how APIs work, but you don't need prior Cloudflare experience. Here's the full checklist:

  • A free Cloudflare account — sign up at dash.cloudflare.com. No credit card is required for the free tier.
  • Node.js 18 or newer — installed on your machine. Download from nodejs.org. You can verify with node --version.
  • Basic JavaScript familiarity — enough to read and paste code. You won't be writing complex logic from scratch.
  • A terminal or command prompt — on Windows, that's CMD or PowerShell; on macOS/Linux, any shell works.
  • Your FAQ content ready — a short list of 3–10 question-and-answer pairs about your product or service.
  • About an hour of focused time — the actual work is about 30 minutes of setup plus 15 minutes of testing, but having a full hour leaves room to debug if something unexpected comes up.
💡 New to Cloudflare? Cloudflare's official Workers Quickstart guide and Vectorize documentation are excellent references if you want to dive deeper after this tutorial.

Why Build Your Own Chatbot Instead of Paying for Intercom or Chatbase?

I'm not against SaaS chatbot tools — if you're running a support team of ten agents, a platform like Intercom genuinely earns its price with ticketing, live handoff, and analytics baked in. But for the majority of website owners who just want visitors to get quick, accurate answers to common questions, you're paying for a lot of machinery you'll never touch. Here's how the free, self-hosted route actually compares:

Option Typical Cost Setup Effort You Own the Data
Intercom Fin AI ~$0.99 per resolved chat + seat fees Low (dashboard setup) No
Chatbase / Tidio (free tier) Free up to a low message cap, then $20–$100+/mo Very low Partially
Cloudflare Workers (this guide) $0 for most low-to-mid traffic sites Moderate (one-time build) Yes, fully

The trade-off is honest: you're spending an hour of setup time in exchange for full ownership and zero recurring cost. If you already understand the basics of how APIs work, you're more than qualified to follow along — you don't need prior experience with Cloudflare specifically.

What You're Building: The Architecture in Plain Terms

Before touching a terminal, it helps to see the moving parts. The whole project is two files talking to four Cloudflare services:

  • A backend Worker — a small serverless function that receives chat messages, looks up relevant FAQ context, calls the AI model, and streams the reply back.
  • A frontend widget script — a single embeddable JavaScript file that draws the chat bubble and window on any page it's loaded into.
  • Workers AI — runs the actual language model (we'll use Meta's Llama 3) at the edge, so you're not paying OpenAI per token.
  • Vectorize — a vector database. It stores your FAQ as numerical "embeddings" so the bot can find the most relevant answer by meaning, not just exact keyword matches. This is the RAG part.
  • KV (Key-Value store) — Cloudflare's globally distributed storage, used here to remember each visitor's conversation between page loads.

None of these run on a server you manage. Cloudflare's edge network (300+ locations worldwide) runs your code close to whoever's visiting the site, which is also why response times stay snappy even on the free tier.

Step 1: Create the Project and Cloudflare Resources

You'll need a free Cloudflare account and Node.js 18 or newer installed. Start by scaffolding a Worker project:

Terminal

npm create cloudflare@latest support-widget
cd support-widget
npm install --save-dev tailwindcss wrangler

Pick javascript when prompted, and choose "no" when asked to deploy immediately — we'll deploy once everything's ready. Next, install Wrangler globally (Cloudflare's CLI) and log in:

Terminal

npm install -g wrangler
wrangler login

Now create the two Cloudflare resources the project depends on — a Vectorize index for your FAQ, and a KV namespace for chat sessions:

Terminal

npx wrangler vectorize create support-faq --dimensions=768 --metric=cosine
npx wrangler kv namespace create CHAT_HISTORY

Keep the ID printed by the second command handy — you'll paste it into the config file in the next step.

Step 2: Wire Everything Together in wrangler.jsonc

This config file tells your Worker which Cloudflare services it's allowed to use. Create wrangler.jsonc in your project root:

wrangler.jsonc

{
  "name": "support-widget",
  "main": "src/worker.js",
  "compatibility_date": "2026-01-01",
  "assets": { "directory": "./public", "binding": "ASSETS" },
  "ai": { "binding": "AI" },
  "vectorize": [
    { "binding": "FAQ_INDEX", "index_name": "support-faq" }
  ],
  "kv_namespaces": [
    { "binding": "CHAT_HISTORY", "id": "PASTE_YOUR_KV_ID_HERE" }
  ]
}
💡 What each binding does: ASSETS serves your static widget files, AI gives the Worker access to run models, FAQ_INDEX connects to your vector database, and CHAT_HISTORY is where visitor conversations get stored.

Step 3: Build the Backend Worker

This is where the actual logic lives — receiving a message, pulling relevant FAQ context, calling the model, and streaming the answer back while saving it to KV. Create src/worker.js:

src/worker.js

const SYSTEM_PROMPT = "You are a concise, friendly support assistant. Use the provided FAQ context when relevant. If you're unsure, say so instead of guessing.";
const SESSION_TTL = 60 * 60 * 24 * 14; // 14 days

// The widget almost always lives on a different domain than the site embedding it,
// so we reflect the real request origin instead of using a wildcard — a wildcard
// origin cannot be combined with credentialed (cookie-based) requests.
function corsHeaders(request) {
  return {
    "Access-Control-Allow-Origin": request.headers.get("Origin") || "*",
    "Access-Control-Allow-Credentials": "true",
    "Vary": "Origin",
    "X-Content-Type-Options": "nosniff",
  };
}

function getSessionId(request) {
  const match = request.headers.get("Cookie")?.match(/widget_sid=([^;]+)/);
  return match ? match[1] : null;
}

async function findRelevantFaq(env, question) {
  const embedded = await env.AI.run("@cf/baai/bge-base-en-v1.5", { text: [question] });
  if (!embedded?.data?.length) return "";
  const matches = await env.FAQ_INDEX.query(embedded.data[0], { topK: 3, returnMetadata: true });
  return matches.matches
    .filter((m) => m.score > 0.6) // drop weak matches instead of feeding the model noise
    .map((m) => `Q: ${m.metadata?.question}\nA: ${m.metadata?.answer}`)
    .join("\n\n");
}

// Workers AI streams Server-Sent Events, but chunk boundaries don't respect line
// boundaries — a single "data: {...}" line can arrive split across two chunks.
// Buffering and only parsing complete lines avoids silently dropped tokens.
function createSseLineBuffer(onEvent) {
  let buffer = "";
  return {
    push(text) {
      buffer += text;
      const lines = buffer.split("\n");
      buffer = lines.pop();
      for (const line of lines) {
        if (line.startsWith("data: ") && !line.includes("[DONE]")) {
          try { onEvent(JSON.parse(line.slice(6))); } catch (_) {}
        }
      }
    },
  };
}

async function handleChat(request, env) {
  const cors = corsHeaders(request);
  const { message } = await request.json();
  if (!message?.trim()) {
    return new Response(JSON.stringify({ error: "Message is required" }), { status: 400, headers: { "Content-Type": "application/json", ...cors } });
  }

  let sessionId = getSessionId(request);
  let session = sessionId ? await env.CHAT_HISTORY.get(sessionId, "json") : null;

  // needsCookie must track whether we actually generated a NEW session id —
  // not just whether a cookie was present. If the cookie survived but the KV
  // entry expired, the old code below would silently start a fresh session
  // without ever re-issuing a cookie, breaking memory permanently.
  let needsCookie = false;
  if (!session) {
    sessionId = crypto.randomUUID();
    session = { history: [] };
    needsCookie = true;
  }

  session.history.push({ role: "user", content: message.trim() });
  const faqContext = await findRelevantFaq(env, message);
  const promptMessages = [
    { role: "system", content: SYSTEM_PROMPT + (faqContext ? `\n\nRelevant FAQ:\n${faqContext}` : "") },
    ...session.history.slice(-8),
  ];

  const modelStream = await env.AI.run("@cf/meta/llama-3-8b-instruct", { messages: promptMessages, stream: true });
  let fullReply = "";
  const decoder = new TextDecoder();
  const lineBuffer = createSseLineBuffer((chunk) => { fullReply += chunk.response || ""; });

  const { readable, writable } = new TransformStream({
    transform(chunk, controller) {
      lineBuffer.push(decoder.decode(chunk, { stream: true }));
      controller.enqueue(chunk);
    },
    async flush() {
      if (fullReply) {
        session.history.push({ role: "assistant", content: fullReply });
        await env.CHAT_HISTORY.put(sessionId, JSON.stringify(session), { expirationTtl: SESSION_TTL });
      }
    },
  });

  modelStream.pipeTo(writable).catch((err) => console.error("Stream error:", err));

  const headers = { "Content-Type": "text/event-stream", ...cors };
  // SameSite=None + Secure is required (not Lax) because this is a cross-site
  // request from the host page's origin to the Worker's origin. Lax cookies
  // are withheld on cross-site fetch calls, which silently breaks memory.
  if (needsCookie) headers["Set-Cookie"] = `widget_sid=${sessionId}; Path=/; HttpOnly; SameSite=None; Secure; Max-Age=${SESSION_TTL}`;
  return new Response(readable, { headers });
}

async function handleHistory(request, env) {
  const cors = corsHeaders(request);
  const sessionId = getSessionId(request);
  const session = sessionId ? await env.CHAT_HISTORY.get(sessionId, "json") : null;
  return new Response(JSON.stringify({ history: session?.history || [] }), { headers: { "Content-Type": "application/json", ...cors } });
}

async function handleSeed(request, env) {
  const cors = corsHeaders(request);
  const providedKey = request.headers.get("X-Seed-Key");
  if (!env.SEED_KEY || providedKey !== env.SEED_KEY) {
    return new Response(JSON.stringify({ error: "Unauthorized" }), { status: 401, headers: { "Content-Type": "application/json", ...cors } });
  }

  const { faq } = await request.json(); // faq: [{ question, answer }, ...]
  if (!Array.isArray(faq) || !faq.length) {
    return new Response(JSON.stringify({ error: "Provide a non-empty faq array" }), { status: 400, headers: { "Content-Type": "application/json", ...cors } });
  }

  const vectors = [];
  for (const entry of faq) {
    const embedded = await env.AI.run("@cf/baai/bge-base-en-v1.5", { text: [entry.question] });
    vectors.push({
      id: crypto.randomUUID(),
      values: embedded.data[0],
      metadata: { question: entry.question, answer: entry.answer },
    });
  }

  await env.FAQ_INDEX.upsert(vectors);
  return new Response(JSON.stringify({ seeded: vectors.length }), { headers: { "Content-Type": "application/json", ...cors } });
}

export default {
  async fetch(request, env) {
    const path = new URL(request.url).pathname;
    const cors = corsHeaders(request);

    if (request.method === "OPTIONS") {
      return new Response(null, { headers: { ...cors, "Access-Control-Allow-Methods": "GET,POST,OPTIONS", "Access-Control-Allow-Headers": "Content-Type, X-Seed-Key" } });
    }
    if (path === "/api/chat" && request.method === "POST") return handleChat(request, env);
    if (path === "/api/history") return handleHistory(request, env);
    if (path === "/api/seed" && request.method === "POST") return handleSeed(request, env);
    return env.ASSETS.fetch(request);
  },
};
⚠️ The bug that breaks memory silently: A wildcard Access-Control-Allow-Origin: * combined with a credentialed (cookie-carrying) request is rejected by every modern browser — and since your widget's domain is almost never the same as the site embedding it, this is a cross-site request by default. On top of that, a SameSite=Lax cookie is withheld on cross-site fetch calls entirely. Get either of these wrong and the chatbot will appear to work in your first message, then quietly forget everything on the next one — no error, no console warning, just a bot with amnesia. That's why the code above reflects the real request origin and uses SameSite=None; Secure.

A few more details worth calling out. The session cookie is HttpOnly, which stops client-side scripts from reading it — a small but important security habit for anything that tracks a visitor across requests. The FAQ lookup filters out weak matches below a similarity score of 0.6, so a completely unrelated question doesn't get padded with irrelevant "context" that confuses the model. And the SSE parser buffers partial lines across chunk boundaries instead of parsing each chunk in isolation — without that buffer, streamed responses can drop random words whenever a JSON line happens to split across two network packets, which is the kind of bug that only shows up intermittently and is miserable to debug after the fact.

Finally, notice that needsCookie is set whenever a new session is created — not just when the incoming request had no cookie at all. If a visitor's cookie survives but their KV entry has expired (after the 14-day TTL), the original approach would silently start a fresh session without ever re-issuing a cookie, leaving that visitor's history unrecoverable on every future visit. It's the kind of edge case that passes a quick manual test and only breaks two weeks into production.

Step 4: Set Up Tailwind CSS for the Widget

Since this widget gets embedded on someone else's site, using a full CSS framework with scoped utility classes avoids clashing with the host page's existing styles. Create tailwind.config.js:

tailwind.config.js

module.exports = {
  content: ["./public/**/*.{html,js}"],
  darkMode: "class",
  theme: { extend: {} },
};

Then add a source stylesheet at src/styles.css with the three standard Tailwind directives:

src/styles.css

@tailwind base;
@tailwind components;
@tailwind utilities;

Then add these scripts to package.json:

package.json (scripts section)

"scripts": {
  "build:css": "npx tailwindcss -i ./src/styles.css -o ./public/widget.css --minify",
  "dev": "npm run build:css && wrangler dev",
  "deploy": "npm run build:css && wrangler deploy"
}

Step 5: Build the Embeddable Widget Script

This is the file that any website will load with a single <script> tag. It builds the chat bubble, handles opening and closing the window, and streams the AI's response into the page. Create public/widget.js:

public/widget.js

(function () {
  const config = {
    baseUrl: window.SUPPORT_WIDGET_URL || "",
    title: window.SUPPORT_WIDGET_TITLE || "Support",
    greeting: window.SUPPORT_WIDGET_GREETING || "Hi! What can I help you with?",
  };
  let messages = [];
  let isOpen = false;
  let isSending = false;

  // Message content comes from the visitor AND from the model — neither is
  // trusted. Without escaping, typing "<img src=x onerror=alert(1)>" into the
  // chat box would execute arbitrary script on the host page.
  function escapeHtml(str) {
    return str
      .replace(/&/g, "&")
      .replace(/
/g, ">") .replace(/"/g, """) .replace(/'/g, "'"); } function injectMarkup() { const styleLink = document.createElement("link"); styleLink.rel = "stylesheet"; styleLink.href = config.baseUrl + "/widget.css"; document.head.appendChild(styleLink); const root = document.createElement("div"); root.id = "support-widget-root"; root.innerHTML = ` <button id="sw-toggle" class="fixed bottom-6 right-6 w-14 h-14 bg-emerald-600 rounded-full shadow-xl z-[99999] text-2xl">💬</button> <div id="sw-panel" class="fixed bottom-24 right-6 w-96 max-w-[calc(100vw-2rem)] h-[560px] max-h-[75vh] rounded-2xl shadow-2xl bg-white hidden flex-col z-[99999]"> <div class="flex items-center justify-between px-4 py-3 border-b font-semibold"> <span>${escapeHtml(config.title)}</span> <button id="sw-close" class="text-gray-400 hover:text-gray-700 text-lg leading-none">✕</button> </div> <div id="sw-messages" class="flex-1 overflow-y-auto p-4 space-y-3"></div> <form id="sw-form" class="flex gap-2 p-3 border-t"> <input id="sw-input" class="flex-1 border rounded-full px-4 py-2 text-sm" placeholder="Type a message..." autocomplete="off" /> <button type="submit" class="px-4 bg-emerald-600 text-white rounded-full text-sm">Send</button> </form> </div>`; document.body.appendChild(root); } function renderMessages(showTyping) { const container = document.getElementById("sw-messages"); const bubbles = messages .map((m) => `<div class="text-sm ${m.role === "user" ? "text-right" : "text-left"}"> <span class="inline-block px-3 py-2 rounded-xl ${m.role === "user" ? "bg-emerald-600 text-white" : "bg-gray-100"}">${escapeHtml(m.content)}</span> </div>`) .join(""); const typingBubble = showTyping ? `<div class="text-left"><span class="inline-block px-3 py-2 rounded-xl bg-gray-100 text-sm text-gray-400">Typing…</span></div>` : ""; container.innerHTML = bubbles + typingBubble; container.scrollTop = container.scrollHeight; } // Same buffering fix as the backend — a chunk boundary can split a single // "data: {...}" line in two, so we hold back any incomplete trailing line. function createSseLineBuffer(onEvent) { let buffer = ""; return { push(text) { buffer += text; const lines = buffer.split("\n"); buffer = lines.pop(); for (const line of lines) { if (line.startsWith("data: ") && !line.includes("[DONE]")) { try { onEvent(JSON.parse(line.slice(6))); } catch (_) {} } } }, }; } async function sendMessage(text) { if (isSending) return; // prevents overlapping requests if someone spams Enter isSending = true; messages.push({ role: "user", content: text }); renderMessages(true); try { const response = await fetch(config.baseUrl + "/api/chat", { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ message: text }), credentials: "include", }); if (!response.ok) throw new Error("Request failed: " + response.status); const reader = response.body.getReader(); const decoder = new TextDecoder(); let reply = ""; let index = null; const lineBuffer = createSseLineBuffer((chunk) => { if (!chunk.response) return; reply += chunk.response; if (index === null) { messages.push({ role: "assistant", content: reply }); index = messages.length - 1; } else { messages[index].content = reply; } renderMessages(false); }); while (true) { const { done, value } = await reader.read(); if (done) break; lineBuffer.push(decoder.decode(value, { stream: true })); } if (index === null) { messages.push({ role: "assistant", content: "Sorry, I didn't catch that — could you try again?" }); renderMessages(false); } } catch (err) { messages.push({ role: "assistant", content: "Something went wrong reaching support. Please try again shortly." }); renderMessages(false); } finally { isSending = false; } } function togglePanel(forceOpen) { isOpen = typeof forceOpen === "boolean" ? forceOpen : !isOpen; const panel = document.getElementById("sw-panel"); panel.classList.toggle("hidden", !isOpen); panel.classList.toggle("flex", isOpen); if (isOpen && !messages.length) { messages.push({ role: "assistant", content: config.greeting }); renderMessages(false); } } function bindEvents() { document.getElementById("sw-toggle").onclick = () => togglePanel(); document.getElementById("sw-close").onclick = () => togglePanel(false); document.getElementById("sw-form").onsubmit = (e) => { e.preventDefault(); const input = document.getElementById("sw-input"); const text = input.value.trim(); if (!text) return; input.value = ""; sendMessage(text); }; } async function loadHistory() { try { const res = await fetch(config.baseUrl + "/api/history", { credentials: "include" }); if (res.ok) { const data = await res.json(); if (data.history?.length) { messages = data.history; } } } catch (_) {} } async function init() { injectMarkup(); await loadHistory(); bindEvents(); } document.readyState === "loading" ? document.addEventListener("DOMContentLoaded", init) : init(); })();
⚠️ Never skip HTML-escaping chat content: Both the visitor's typed message and the model's reply get inserted into the page with innerHTML. If you insert that text raw, a visitor typing something like an <img onerror=...> tag into the chat box gets executed as real HTML on whatever site the widget is embedded on — a stored cross-site scripting vulnerability. The escapeHtml() helper above is not optional polish; skip it and you've shipped a security hole on every site that installs your widget.

The whole thing is wrapped in an IIFE (a function that runs immediately and doesn't leak variables into the host page's global scope) — an important habit for any script meant to be dropped onto someone else's site. The streaming loop in sendMessage() updates the same message bubble in place as tokens arrive, rather than re-rendering the whole chat window on every chunk — that's what produces the smooth "typing" effect instead of visible flicker. The try/catch around the fetch call means a dropped connection shows the visitor a friendly fallback message instead of a chat window that just hangs forever with no feedback — something the bare-bones version of this pattern almost always skips.

Step 6: Test Locally, Then Deploy

One file is still missing: since env.ASSETS.fetch() serves whatever's in ./public, visiting the root URL right now would just 404 — there's no page there yet. Add a minimal demo page so wrangler dev actually has something to show you. Create public/index.html:

public/index.html

<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8" />
  <title>Support Widget — Local Demo</title>
</head>
<body style="font-family: sans-serif; padding: 40px; max-width: 640px; margin: 0 auto;">
  <h1>Local demo page</h1>
  <p>This page only exists so the dev server has something to render at the root.
    The chat bubble should appear in the bottom-right corner — click it to test.</p>
  <script>
    window.SUPPORT_WIDGET_URL = ""; // same-origin while testing locally
    window.SUPPORT_WIDGET_TITLE = "Ask Us Anything";
  </script>
  <script src="/widget.js"></script>
</body>
</html>

Run the dev server and open it in your browser:

Terminal

npm run dev
# opens at http://localhost:8787

Click the chat bubble in the corner and send a test message — you should see it stream back token by token. If nothing happens, open your browser's console first; a CORS or cookie error will show up there immediately.

⚠️ Local testing note — cookies won't work on plain HTTP: The cookie is set with SameSite=None; Secure, which is required for cross-site requests in production. However, modern browsers reject the Secure flag on plain http://localhost connections. This means the session cookie will not be set during local testing with wrangler dev, and the chatbot will appear to forget the conversation after each message. This is expected — the cookie works correctly once deployed to Cloudflare's HTTPS endpoint. To test locally with full session persistence, use wrangler dev --remote (which uses your live Cloudflare account with HTTPS) or tunnel through a tool like ngrok.

Once you're happy with it locally, deploy with a single command:

Terminal

npm run deploy

You'll get a live URL that looks like support-widget.YOUR-SUBDOMAIN.workers.dev. The bot won't be able to answer FAQ questions yet, though — the Vectorize index is still empty. That's what the next step fixes.

Seed Your FAQ Data into Vectorize

Before the chatbot can answer real questions, you need to convert your FAQ into embeddings and store them in the vector index. The handleSeed() function already added to worker.js (in the code block from Step 3, under the /api/seed route) handles this — it accepts a list of question/answer pairs, embeds each question, and upserts the result into Vectorize. It's protected by a secret key so a stranger can't overwrite your FAQ.

Set that secret before deploying:

Terminal

npx wrangler secret put SEED_KEY
# paste any random string when prompted — you'll reuse it below

Then call the endpoint once with your FAQ content:

Terminal

curl -X POST https://support-widget.YOUR-SUBDOMAIN.workers.dev/api/seed \
  -H "Content-Type: application/json" \
  -H "X-Seed-Key: YOUR_SECRET_HERE" \
  -d '{
    "faq": [
      { "question": "What are your shipping times?", "answer": "Orders ship within 2 business days and arrive in 3-7 days depending on location." },
      { "question": "Do you offer refunds?", "answer": "Yes, full refunds are available within 30 days of purchase." },
      { "question": "How do I contact support?", "answer": "Email support@example.com or use this chat widget." }
    ]
  }'

A successful response returns {"seeded": 3} — confirming three FAQ entries were embedded and stored. Re-run this same call any time your FAQ content changes; there's no need to clear the index first since each entry gets a fresh unique ID.

Step 7: Embed the Widget on Any Website

This is the whole point — adding your chatbot to any site, including one built on a completely different stack, takes two lines before the closing </body> tag:

HTML

<script>
  window.SUPPORT_WIDGET_URL = "https://support-widget.YOUR-SUBDOMAIN.workers.dev";
  window.SUPPORT_WIDGET_TITLE = "Ask Us Anything";
</script>
<script src="https://support-widget.YOUR-SUBDOMAIN.workers.dev/widget.js"></script>

That's it. Because the widget is a self-contained script with its own styles, it won't fight with your site's existing CSS — it'll work the same whether the host page is built with WordPress, a static site generator, or a custom Shopify theme.

5 Mistakes I See People Make With Self-Hosted Chatbots

  1. Skipping the FAQ context entirely. Without RAG, you're just running a generic chat model that knows nothing about your product. It'll hallucinate confidently. Even a handful of well-written Q&A pairs makes a noticeable difference in accuracy.
  2. Not capping conversation history. If you send the entire chat history to the model on every message, both your latency and your neuron usage climb fast. Trim to the last 6–10 messages, which is plenty of context for support-style conversations.
  3. Using a wildcard CORS origin with cookies. Your widget is hosted on a different domain than the page it's embedded in, and a wildcard Access-Control-Allow-Origin: * simply doesn't work once cookies are involved — browsers block it outright. We covered the fix (reflecting the real origin, plus SameSite=None) back in Step 3, but it's worth repeating because it's the single most common reason a self-hosted chatbot "loses" the conversation after the first message.
  4. No fallback for "I don't know." Bake an instruction into your system prompt that tells the model to admit uncertainty rather than invent an answer. A wrong answer damages trust far more than an honest "let me connect you with a human."
  5. Never re-seeding Vectorize after FAQ updates. Your live chatbot only knows what's in the vector index — if you update your FAQ page but forget to re-embed the new content, the bot keeps answering from stale data indefinitely.

Cloudflare's Free Tier: What You Actually Get for $0

Before you commit to this route, it's worth knowing exactly where the free tier ends. From my own testing, these limits are generous enough for the vast majority of small-to-medium sites:

  • Workers: 100,000 requests per day
  • Workers AI: 10,000 "neurons" per day (roughly a few hundred short conversations, depending on message length)
  • Vectorize: 5 million vector query/insert operations per month
  • KV: 100,000 reads and 1,000 writes per day

If you outgrow these — congratulations, your site has real traffic — Cloudflare's paid tiers scale up without requiring you to rebuild anything, since you're already on their infrastructure.

Where to Take This Project Next

Once the basic widget is live, a few upgrades are worth exploring. You could swap Llama 3 for a different open model to compare response quality — our Qwen vs GPT vs Gemini comparison is a good starting point for understanding the trade-offs between models. If you'd rather orchestrate the chatbot's logic visually instead of hand-coding it, our n8n automation guide covers a no-code alternative worth comparing against this hand-built approach. And if the difference between an "AI agent" and a simple chatbot has you curious, we cover that distinction in our guide on building a free team of AI agents for your website.

Want to version-control your Worker code and collaborate with others on it? Our Git and GitHub beginner's guide covers exactly that, and it pairs well with pushing this project to a repository before you start customizing it further.

Troubleshooting Common Issues

Even with correct code, things can go wrong. Here are the most common issues and how to fix them:

❌ Chat works for one message, then "forgets" the conversation

Cause: The cookie either wasn't set (check DevTools → Application → Cookies) or has the wrong SameSite attribute. This is the #1 most common issue.
Fix: Ensure you're on HTTPS when testing (use wrangler dev --remote instead of plain wrangler dev), and verify your browser isn't blocking third-party cookies.

❌ CORS error in the browser console

Cause: The Access-Control-Allow-Origin header is missing or is a wildcard * while credentials: "include" is used.
Fix: The corsHeaders() function in the Worker already handles this — make sure it's being applied to every response. If you're testing locally, try deploying first since localhost has its own CORS quirks.

❌ 404 when loading widget.js or widget.css

Cause: The ASSETS binding in wrangler.jsonc points to ./public, but the files aren't there or the path is wrong.
Fix: Run npm run build:css first to generate widget.css, and confirm both widget.js and widget.css exist in the public/ directory.

❌ "Model not found" error from Workers AI

Cause: The model name @cf/meta/llama-3-8b-instruct may have been deprecated or renamed. Cloudflare periodically updates their model catalog.
Fix: Check the Workers AI model catalog for the latest available models, and update the name in env.AI.run(). Popular alternatives include @cf/meta/llama-3.1-8b-instruct or @cf/mistral/mistral-7b-instruct-v0.1.

Customizing the Widget's Look and Feel

The widget uses Tailwind utility classes, which makes customization straightforward. Here are a few common changes:

  • Change the accent color — Replace every bg-emerald-600 in widget.js with your brand color, e.g. bg-blue-600. The same applies to text-emerald-600 and the green gradient in the border.
  • Move the chat bubble — Adjust the bottom-6 right-6 classes. For a left-aligned widget, use bottom-6 left-6 instead.
  • Change the widget size — Modify w-96 (width) and h-[560px] (height) in the panel div.
  • Replace the bubble icon — Change the 💬 emoji in the toggle button to your own icon or SVG.
  • Add your own CSS overrides — Add styles to src/styles.css before the Tailwind directives, or add a second stylesheet after widget.css.

Tracking Chatbot Usage and Performance

Once your widget is live, you'll want to know how it's being used. Cloudflare's dashboard gives you basic metrics (request counts, error rates), but for deeper insight, consider these approaches:

  • Cloudflare Analytics dashboard — Visit the Cloudflare Dashboard → Workers & Pages → your Worker. You'll see request volume, CPU time, and error rates out of the box — no extra code needed.
  • Log into KV events — Add a small logging step inside handleChat() that writes anonymized usage data (question asked, whether a FAQ match was found, response time) to a separate KV namespace. This lets you analyze which questions are most common and whether the bot is answering them well.
  • Manual review via wrangler tail — Run npx wrangler tail to watch live logs from your deployed Worker. Useful for debugging real-time issues.
  • User feedback integration — Add a "Was this helpful?" thumbs-up/thumbs-down button after each bot response. Store the feedback in KV to measure satisfaction over time.

Going to Production: What's Missing

The widget as built above works, but a production deployment should add a few more layers:

  • Custom domain — Instead of the *.workers.dev subdomain, point your own domain (e.g. widget.yourdomain.com) to the Worker. Add a routes key in wrangler.jsonc: "routes": [{ "pattern": "widget.yourdomain.com", "custom_domain": true }].
  • Rate limiting — Add basic rate limiting to prevent abuse. Cloudflare's Rate Limiting example shows how to add a counter in KV and reject requests after a threshold.
  • Error monitoring — Integrate with Workers Observability to get alerts when error rates spike.
  • Human handoff — For questions the bot can't answer, add a "Talk to a human" button that sends the conversation transcript to your support email or a Slack webhook.

Privacy, GDPR, and Compliance Considerations

If your website serves visitors in the EU, UK, or California, a chatbot that stores conversation history needs a few compliance basics:

  • Cookie consent — The widget_sid cookie is a strictly necessary cookie (it powers the conversation feature), but your cookie consent banner should still disclose it. Update your privacy policy to mention the cookie and its 14-day retention period.
  • Data retention — The 14-day TTL on KV entries means conversations are automatically deleted. If you need longer retention for analytics, anonymize the data and store it separately with a clear retention policy.
  • Right to deletion — Add a DELETE /api/history endpoint that clears the visitor's KV entry, allowing users to exercise their right to be forgotten under GDPR Article 17.
  • Data processing addendum — If you're using this widget for a client site, Cloudflare acts as a data processor. Your clients may need a DPA (Data Processing Addendum) with Cloudflare, which is available in your Cloudflare dashboard under Account → Agreements.

What Happens When You Outgrow the Free Tier

Cloudflare's free tier is generous, but it's worth planning for growth. Here's what the paid tiers look like and when you'd need them:

Service Free Limit Paid Starts At When You'd Need It
Workers 100k req/day ~$5/month (Workers Paid) ~3,000 visitors/day
Workers AI 10k neurons/day ~$0.001/1k neurons ~500 conversations/day
Vectorize 5M operations/mo Usage-based ~50k FAQ queries/day
KV 100k reads + 1k writes/day ~$0.50/month for 1M reads ~3k visitors/day

Even the Workers Paid plan at $5/month is dramatically cheaper than any SaaS chatbot for similar traffic levels. The key advantage is that pricing is usage-based — you only pay for what you actually use, and your existing code works without modification.

Final Thoughts

What you've built here isn't a toy — it's a legitimate, production-capable alternative to tools that routinely charge hundreds of dollars a month. The pieces (a serverless function, a vector database, and a key-value store) are the same building blocks powering much larger AI products; you're just running them at a scale that happens to fit inside Cloudflare's free tier.

The parts most tutorials skip — trimming conversation history, handling CORS properly, giving the model permission to say "I don't know" — are exactly the details that separate a demo from something you'd actually trust on a live website. Get those right, and the difference between this and a $99/month SaaS widget mostly comes down to who owns the setup script.

If you build this out, I'd genuinely like to know how the free-tier limits hold up on your traffic — drop a comment once it's live.

📬

Liked the hands-on approach?

Join hundreds of subscribers and get practical AI, security, and networking guides — projects, not theory — delivered to your inbox.

Yes, Subscribe Me! ✉️

🔒 No spam, ever. We respect your inbox.

Frequently Asked Questions

❓ Is a Cloudflare Workers AI chatbot really free?

For most low-to-mid traffic sites, yes. The free tier includes 100,000 Worker requests, 10,000 AI neurons, 5 million Vectorize operations, and 100,000 KV reads per day. You'd need real, sustained traffic before hitting any of those ceilings, and even then, Cloudflare's paid tiers are usage-based rather than flat per-resolution fees.

❓ Do I need to know JavaScript to follow this guide?

Basic familiarity helps, but you don't need to be an experienced developer. Most of the work is copying configuration and running CLI commands. The code itself is short enough that reading through the explanations alongside it should get you unstuck if something doesn't work as expected.

❓ What is RAG, and why does this chatbot need it?

RAG (Retrieval Augmented Generation) means the chatbot looks up relevant facts from your own FAQ before generating an answer, instead of relying purely on what the underlying language model already "knows." Without it, the bot has no idea about your specific product, policies, or pricing — it can only speak in generalities.

❓ Can I use a different AI model instead of Llama 3?

Yes. Workers AI hosts several open models beyond Llama 3, and you can also route requests to external providers like OpenAI or Gemini if you're willing to pay per-token for a different quality tier. Swapping models generally only requires changing the model name in the env.AI.run() call, though prompt formatting can vary slightly between models.

❓ Will this widget slow down my website?

Minimal impact if implemented as shown. The widget script is small, loads asynchronously, and doesn't block the rest of the page from rendering. The heavier work — running the AI model and searching Vectorize — happens entirely on Cloudflare's servers, not in the visitor's browser.

❓ How do I keep the chatbot's FAQ answers up to date?

Whenever your FAQ content changes, re-run your embedding/upsert script against Vectorize so the new question-and-answer pairs are searchable. The chatbot only knows what's stored in the vector index — editing your website's FAQ page alone doesn't update what the bot can retrieve.

📌 Found this guide useful? Share it with someone tired of paying monthly for a chatbot widget, and explore more hands-on AI and automation guides at Valley4Techs.

Add Valley4Techs as a Preferred Source

Follow us on Google News for the latest updates

Add Now
Mostafa Amaan
Mostafa Amaan
Technical educational content creator on my blog and YouTube channel. My goal with this content is to eradicate information technology literacy.
Comments