Analytics & TestingArtificial IntelligenceContent MarketingEmail Marketing & Automation

Ask Martech v3.0: How I Built My Own Chatbot on WordPress Using Cloudflare

Martech Zone grew large enough to attract bots that were looking for nothing in particular. They fired random terms at WordPress search, a query string at the end of the address, and each request made the database scan the archive. The admin timed out. Server resources kept spiking, and my site health was a rollercoaster. I wrote about this (and the solution) in How to Prevent Resource Exhaustion Attacks on WordPress Search.

I decided to disable search altogether and replaced that public query-string search with a chatbot. A box that accepts any string and runs it across every post is an invitation. The reply uses the content of the site to put context around the question: what this archive actually says, which page said it, and what it leaves unsaid.

Why didn’t I just use a third-party chatbot? This is a key question and there are a few reasons. Primarily, by having the chatbot self-hosted, I removed all of the terrible CWV impact that most chatbots incur. My chatbot doesn’t impact site speed whatsoever. Second, I have some custom post types and unique data associated with my site, like acronyms, a conference calendar, and retail holiday calendar. Those required some unique processes.

I started this with a very rudimentary chatbot built using Cloudflare services after reading how Joost de Valk built his. My first results weren’t that great, but it was a start. Now, quite a few versions later, users can ask questions like How do I set up DMARC? The chatbot responds in detail.

How Ask Martech Reached v3.0

Ask Martech started as that replacement. The first useful version was retrieval-augmented generation (RAG). Passages from the archive were embedded, and a model drafted only from what came back. The next section says what those two actions actually are.

Retrieval had to get stricter before the writer got larger. At one point, a shared tag was enough to let the newest post outrank the page that defined the term. Calendar questions needed their own sense of which month was actually being asked about. The writer then moved to gpt-oss-120b, which can stay with a longer excerpt and still answer briefly.

The remaining mistakes weren’t missing pages. They were sentences that combined two articles or gave a first-person line to the person the visitor had named. Clef-flash now checks a page before the draft, and checks a sentence after it. AI Gateway sits on every model call, so it logs the calls and caches identical requests. The window you open today is labeled v3.0.

WordPress and Cloudflare

WordPress still holds the library and the chat. My custom plugin was largely written with Claude first, and now Grok. It queues the index, renders the widget, and signs the requests to the edge. Cloudflare Workers run the retrieval and call the hosted models. The question is answered from this archive.

Two programs do that work. The WordPress plugin keeps the articles, the widget, and the settings. The Cloudflare Worker retrieves the passages, judges them, and writes the reply. A question goes from the browser straight to the Worker. A publish reaches the Worker only after cron, and only with a signature.

WordPress Cloudflare Architecture

What Retrieval-Augmented Generation Actually Does

A chat model on its own answers from patterns it learned in training. It has no window into the articles on Martech Zone, and a smooth sentence is easy for it to invent. Retrieval-augmented generation changes the order of the work. First the system retrieves passages from this archive. Then the model writes a reply from those passages. The archive stays the source. The model is the writer.

Picture a closed-book assignment. You hand someone a few highlighted printouts and a question, and you tell them to answer from the printouts. If the printouts never mention the step, they say so. They do not fill the gap from memory. That is the whole design. The machinery below is how the printouts are chosen, and how a second reader checks the sentences after they are written.

Embedded is the part that sounds more mysterious than it is. A passage goes to an embedding model, and the model returns a long row of numbers. You cannot read the numbers as a summary. They are a position. Two passages about setting up email authentication sit near each other. A passage about a conference sits somewhere else. The question is converted with the same model, and the search is a nearest-neighbor lookup: which stored rows are closest to the question’s row?

The words are stored beside those numbers. Finding a nearby row only tells the system which passage to open. The writer later reads the words. It never sees the row of numbers.

How a Published Article Joins the Library

When a post is published or updated, the plugin does not call out during the save. It queues a signed request, and the editor returns immediately. A cron run delivers that request, usually within about ten minutes. Unpublishing removes that URL from the index.

The library lags a new article by that cron gap. A question asked in the gap still uses the previous passages for that address. After the run, those passages are replaced with the new ones.

An article is too long to store as one position. The Worker cuts it into passages of a few paragraphs, about 500 tokens, with a short overlap. A token is a piece of a word, so a passage is a section, not the whole page. The overlap keeps a step that lands on the cut inside the next passage as well, so the setup instructions are less likely to be sliced in half and lost.

Each passage is embedded with bge-base-en-v1.5 on Workers AI. That is the same model used later on the question, which is required. Two different embedding models would draw two different maps, and closest would be meaningless. While the numbers are calculated, the page title and its topics are placed in front of the passage. A short question can then land near a page whose title names the subject, even when the matching words are a small part of the body. The passage stored for the writer does not include that prefix. It is the article text.

The rows of numbers go to Vectorize. The passage text, the titles, and the topic keys go to Workers KV, a key-value store (KV). Read those as two shelves. Vectorize is the map of positions. Workers KV is the cabinet that holds the words. A lookup starts on the map and finishes by pulling the text from the cabinet.

Conferences and retail holidays are their own content, separate from ordinary posts. Dates are read in America/Indiana/Indianapolis, the site timezone. A question about this month is answered against that clock, not against UTC, which would already be tomorrow on many evenings here.

How One Question Moves Through the System

Open Ask Martech Zone and type “How do I set up DMARC?” The widget sends that sentence to a Cloudflare Worker and waits for a finished answer. WordPress search does not run. Nothing scans the database for those words.

These are the requests, in the order they happen.

WordPress Cloudflare Chatbot Logic
  1. The Worker cleans the question. Extra spaces come out, and a very long question is cut off. This one is short, in English, and it names a subject, so it is looked up as written. A short follow-up such as “what about the reports?” can borrow the previous turn, because it has no new subject of its own. Words such as “this month” are rewritten to the real month on the Indianapolis clock before the lookup, so the embedding is not left to guess which month you meant.
  2. bge-base-en-v1.5 turns the question into a row of numbers. It is the same kind of row that was stored for every passage. This request is an embedding call. It does not write an answer.
  3. Vectorize returns the passages whose rows sit closest. Closeness means the same neighborhood. It does not yet mean the page answers the question. On a live run of this DMARC question, the nearest pages included the DMARC report viewer, a deliverability article, pages about other email services, a page about how domain names are looked up, and a page about building an SPF record.
  4. Title and topic signals then adjust that order. A title that matches the question is lifted. A tag can support a page the numbers already found. A shared tag does not give every tagged page one flat score. That flat score used to let the newest tagged post outrank the page that defined the term.
  5. The passage text is loaded from Workers KV. The map found the shelf. The cabinet supplies the words the later steps will read.
  6. Clef-flash reads each retrieved page against the question and returns a probability that the page supports an answer. It still does not write the reply. A page under 0.35 is set aside, unless the title is a strong match or the page is a calendar. If the check fails, returns no probability, or would drop every page, the ranked list stands. A weak page from the archive is safer than an answer drawn from a page the archive does not contain.
  7. The pages that remain are numbered and handed to gpt-oss-120b with your question. The instructions say to use only those excerpts, to call the publication Martech Zone, and to number the steps of a how-to. The model does not open the rest of either article, and it does not open the web. On this DMARC run, the writer received two pages: the DMARC report viewer and the guide to defending your brand from phishing. The neighboring email pages stayed out of the prompt.
  8. Clef-flash reads the draft next. A sentence that cites a source is checked against that source. A name, or any title-case phrase of two or more words, has to appear in the supporting text. If it does not, that sentence is removed and is not restored. When nothing remains, you see one line: Martech Zone does not say that.
  9. Because the question begins with How do I, the reply is treated as steps. Bullets are renumbered. A question that asks what the options are stays a bullet list. The widget draws the marker on the same line as the item’s first word.
  10. You then see the reply, citations such as [1], and the source titles under the reply. The DMARC run came back as a short numbered procedure, with those two pages underneath. Open the citations and compare the steps to the pages. That comparison is the check you have as a reader. If a later change pulls a more direct setup page into the prompt, the steps can change.

While retrieval is being tuned, list mode returns the ranked pages and doesn’t call the writer at all. That shows the neighborhood before Clef-flash and the writer narrow it.

Each service in that path has one job:

ToolJob
VectorizeFinds passages whose meaning sits near the question.
Workers KVHolds the passage text, title keys, topic keys, and ranking settings.
bge-base-en-v1.5Turns pages and questions into rows of numbers on one map.
gpt-oss-120bWrites the reply from the excerpts that remain.
Clef-flashReturns a support probability. It does not write the chat reply.
AI GatewayLogs every model call and caches an identical request for ten minutes.

None of those services replace the archive. They decide which published sentences are allowed into the reply.

Why the Nearest Page Is Not Always the Right Page

The DMARC trace is the reason the early quality work was ranking, not a bigger writer. A flat topic score wiped out the difference between two posts that share a tag. The topic signal now corroborates the vector score, and a tag-only candidate sits lower, as a fallback.

Calendar questions use a separate proximity score so a request for this month prefers this month’s bucket. Change those knobs only after scoring the same questions before and after. The writer never sees that machinery. It only sees the passages that survived it.

What the Writer Is Allowed to Know

By the time the writer runs, its world is your question plus a handful of numbered excerpts. The writer is gpt-oss-120b, an open model from OpenAI, running on Workers AI. An earlier instruct model could draft from excerpts. This one follows a longer passage and still answers briefly. It is an LLM used as a writer, not as a source. Temperature stays low because the job is to state a supported fact, not to invent a smoother one. Any reasoning the model does stays inside the call. The visitor does not see it.

The prompt treats the numbered excerpts as independent articles. A book, a job, or a date in one excerpt is not borrowed for the subject of another. Inside a source, I and we belong to the person who wrote that article. They do not belong to the person, company, or product named in the question. When the reply names this publication, it says Martech Zone.

That rule is there because a larger writer fixed fluency and still mixed people up. Two excerpts can sit in the same prompt: I might mention a writer and a personal anecdote about a company or technology I worked with that illustrated the book. Now the logic is inaccurately making a connection to my anecdote and their work. The excerpt never did. Retrieval did not make that mistake. The draft did. Finding a better page would not have stopped it, because both pages were real. The sentence that combined them was not.

The Check Before the Draft and the Check After It

Steps 6 and 8 in the walkthrough are the same tool, run twice. Clef-flash on Workers AI returns a probability that a statement is supported. The writer stays the writer. Clef-flash never replaces it.

Before the draft, one call scores each retrieved page against the question. That is what narrowed the DMARC neighborhood down to the two pages the writer was allowed to see. A page under 0.35 is dropped, unless the title is a strong match or the page is a calendar. If the check fails, returns no probability, or would drop every page, the original ranking stands.

After the draft, a second call scores each cited sentence against the sources it names. A sentence with no citation stays, unless it assigns a book, a job, or an employer. Those claims are checked against every excerpt. A name, or any title-case phrase of two or more words, has to appear in the supporting text. If it does not, that sentence is removed and is not restored. When nothing remains, the visitor sees one line.

Martech Zone does not say that.

Ask Martech, when the excerpts do not name the subject

If the judge rejects every sentence on probability alone, the draft is kept. A list of upcoming holidays should not vanish because a cautious score disliked the format. A leading yes stays attached to the sentence it introduces, so the yes cannot remain after that sentence is removed.

Where Every Model Call Is Logged

The embedding, the writer, both Clef-flash checks, and translation all go through AI Gateway, under the ID martech-ask. The gateway logs calls and caches identical requests for ten minutes. Ask the same question twice in that window and the second reply can come from the cache. The prompt includes today’s date, so a cached answer about what is upcoming cannot still be upcoming tomorrow. Create the gateway in the Cloudflare dashboard. The credential that deploys the Worker is not allowed to create one.

What You See in the Window

The header reads Ask Martech Zone, with a small v3.0 beside the title. A citation such as [1] becomes a link to that source, and the source titles sit under the reply. For the DMARC question, those titles are the pages the writer was given. Procedures stay numbered, including when the writer emits dashes. A question that asks what the options are stays a bullet list, including when the writer emits numbers. The marker and the first words of each item share a line. Code in the dark theme stays light on a light chip.

The Plugin Settings Pages

In WordPress admin, the menu is Ask Martech. Configuration stores the Worker URL, the shared secret, and which post types are indexed. Index Management holds rank tuning and the calendar refresh. Chat Widget holds the title, the colors, and the avatar. Point the plugin at the Worker on Configuration before the first index run.

Ask Martech Cloudflare WordPress Chatbot Weights and Index Management

What a Question Still Cannot Do

The answer comes from the indexed archive. There is no live web search beside it. The bot cannot credit an author the passage never names. It will refuse a claim rather than borrow one from a neighboring article. That refusal is what the later checks are for. Finding pages came first. Deciding whether a sentence may stand came after. Leave the 0.35 floor where it is until the same questions are scored with the check on and off.

You can ask How do I set up DMARC? yourself and compare the citations with the walkthrough above. As the pages that survive, and the steps in the reply, keep improving, the example in this article will change with them.

Related Articles