# Clanker Support — full blog content
> An AI-powered support agent for your site — install it with one script tag or one React Server Component (@clankersupport/widget-rsc on npm). It answers from your docs and sources, then escalates to your team. Open source (MIT) and self-hostable (bring your own keys); the hosted version has flat monthly plans from $19/mo (50% off right now) with no per-seat fees.
This file concatenates the full text of every clankersupport.com blog post below, newest first. For a short link map of the whole site, see https://clankersupport.com/llms.txt. For the full product documentation in the same format, see https://docs.clankersupport.com/llms-full.txt.
# Why nobody clicks your chat widget
URL: https://clankersupport.com/blog/why-nobody-clicks-your-chat-widget
Published: 2026-08-19
Category: Guides
You installed the widget, checked the dashboard a week later, and found three conversations. Two were you. Here are five small changes that turn a decorative bubble into actual conversations.
You did everything right. Picked a support tool, pasted the script tag, matched the brand color, told the team. A week later you open the dashboard and there are three conversations. Two of them are you, testing it.
The widget isn't broken. It's being ignored, which is worse, because nothing shows up in an error log.
Here's the thing about that little circle in the corner: clicking it is a small action that feels expensive. The visitor has to start a conversation with a stranger, in a blank box, with no idea what happens next or how long it takes. Most people look at that deal and go back to skimming your pricing page with a question still in their head.
Every fix below removes one reason not to click. None of them takes more than a few minutes.
## 1. Stop waiting to be clicked
A bare launcher makes the visitor do all the work: notice it, decide it's worth it, open it, compose a message. That first step is the most expensive one in the whole funnel, and the widget does nothing to help.
Teaser bubbles flip it. A line or two of real text floats up above the closed launcher ("Questions about pricing? I can help"), and suddenly the visitor isn't starting a conversation. They're replying to one that's already open. Replying is cheap. Starting is not.
In Clanker Support, each line of your welcome message becomes its own teaser bubble automatically. Write two short lines and you get a little opening move on every page. If a visitor dismisses them, they stay gone for the session. Nagging is the fastest way to teach people to hate the corner of your website.
## 2. Give people their first sentence
Open any chat widget and you get a cursor blinking in an empty box. Now the visitor has to write. What do I call this thing? Is my question too vague? Am I about to talk to a bot that answers in riddles?
Composing is work, and work loses to the back button.
Starter questions fix this with one tap. Three or four chips under the greeting ("What's included in the free plan?", "Do you integrate with Shopify?") and the visitor's first message is already written. As a bonus, the chips quietly tell them what the agent is actually good at, which sets the conversation up to go well.
Don't invent these. Pull them from your last month of tickets. The questions people ask by email are the questions they'd tap in the widget, if only you'd offer them.
## 3. Rewrite the greeting like a person wrote it
"Hi! How can I help you today?" is what every widget on the internet says. Visitors have seen it a thousand times, so it reads as furniture. Generic greeting, generic expectations, no click.
Compare:
> Hi! How can I help you today?
with:
> Hey, I'm the Acme support agent. Pricing, setup, weird CSV imports: ask me anything. If I get stuck, a human takes over.
The second one does three jobs in two sentences. It says what the agent covers, it sounds like someone actually typed it, and it answers the fear nobody voices out loud: "what if the bot can't help me?" You already know the answer to that objection. Put it in the greeting instead of hoping people find out on their own.
## 4. Ask for less before the conversation
Some widgets open with a form. Name, email, sometimes a dropdown for "topic." Every field is a toll booth between a visitor and their question, and plenty of people just turn around.
We ship with the pre-chat form off. The widget opens straight into the conversation, and you ask for an email only when there's a real reason, like an escalation that needs somewhere to send the reply. You will capture fewer email addresses up front this way. You'll also have far more conversations, and a conversation tells you more about a prospect than an email field they filled with "asdf@asdf.com" anyway.
## 5. Hand off to humans fast, and say so
Most visitors test a widget before they trust it. The first question is a softball. What they're really checking is whether this thing loops them through canned answers forever, because everyone has been trapped in that maze before, and nobody goes back in twice.
So make the exit visible and make it early. We default to offering a human handoff after three exchanges, and honestly, for high-stakes topics like billing you may want it sooner. It feels backwards to advertise the escape hatch from your own bot. It isn't. The visitor who sees a clear path to a person asks the hard question. The one who doesn't ask the easy question, gets a decent answer, and still leaves with the hard one unasked.
We wrote more about scoping the handoff in [our guide to AI first-response layers](https://clankersupport.com/blog/reducing-support-tickets-with-ai-first-response).
## How you'll know it's working
Three numbers, checked weekly, no dashboard archaeology required:
- **Conversations started.** The raw count. This is the number the five fixes above move.
- **Escalation rate.** The share of conversations that reach a human. High isn't bad, and neither is low. What you want is a number that matches your intent: a filter in front of your team, or a full first line of support.
- **Ratings.** Thumbs on individual answers and the end-of-conversation score tell you whether the conversations you earned were worth having.
The bubble in the bottom-right corner of this page has all five of these turned on, so you can judge for yourself how it feels from the visitor's side. Ask it something you'd normally dig through docs for.
And if you'd rather your own corner stopped being decorative: it's [one script tag](https://app.clankersupport.com), and the first 7 days are free.
# Your customers write in 12 languages. Your support agent answers in all of them.
URL: https://clankersupport.com/blog/multilingual-ai-support-agent
Published: 2026-08-11
Author: Omar
Category: Announcements
I spent an hour throwing 12 languages at our support agent — Spanish, Japanese, Arabic, even Moroccan Darija in Arabizi. It answered every one in the customer's language. There's no language setting. And now it does it out loud.
Last night I spent an hour trying to break my own product.
One conversation. Twelve languages. I asked our support agent the same question — free trial, pricing — in Spanish, then French, then Japanese, then Korean, then Arabic. Switching mid-thread. No warning, no pattern, a different alphabet almost every message.
It answered. Every time. In theirs.

Look at the Arabic one. The whole conversation flips right-to-left — the question, the answer, the pricing list. It reads like it was written by someone who grew up writing Arabic.
## There is no language setting
This is my favorite part.
No dropdown. No locale file. No "multilingual add-on" line on your bill. You don't tell it what languages your customers speak — because you don't know. Nobody knows. Your next customer decides that, not your settings page.
Your customer writes. The agent detects. The agent answers. Same knowledge, their language. That's the entire feature.

That's Spanish answered in full — trial, pricing, sources — while the next question is already going in below it. In French. Same thread, no settings touched in between.
My personal stress test was Moroccan Darija, typed the way we actually type it — Latin letters, numbers standing in for sounds no alphabet has. It answered the way my friends text me. If it handles that, it handles your customers anywhere.
## And now it talks
There's a phone icon in the widget now.
Your customer taps it and they're on a live voice call with your support agent. Same knowledge. Same languages. They talk, it talks back — out loud, in real time. (Scale plan.)
I keep calling it just to hear it pick up.
## When it's stuck, you get a human
One more thing, because it's the reason Clanker Support exists at all: when the agent can't answer, it doesn't improvise. It hands the conversation to your team — whole thread attached, nothing for the customer to repeat.
An agent that speaks every language your customers do and still knows when to say "let me get you a person." That's the product.
## Try your language on it
The chat bubble on this page — bottom right — is the live agent. Not a demo build, not a rehearsal. The same one from the video.
Ask it something in your language. Right now. Darija welcome.
And if it wins you over: one script tag, and it's on your site doing this for your customers — in theirs. Free for 7 days.
→ [Start your free trial](https://app.clankersupport.com)
# We diffed our marketing site against our codebase. Six claims didn't survive
URL: https://clankersupport.com/blog/marketing-site-codebase-audit
Published: 2026-07-26
Category: Engineering
We put the homepage in one tab and the repo in the other and checked every falsifiable sentence against the code that would have to make it true. Six claims failed, one page undersold us, and an adversarial re-pass caught overclaims we wrote during the honesty pass itself. Every fix shipped in one public pull request.
Last week we put our marketing site in one tab and our codebase in the other and checked every falsifiable sentence on the site against the code that would have to make it true. Six claims failed the check, and the people most likely to notice were exactly the people we most need to convince.
We build [Clanker Support](https://github.com/theopenco/llmchat), an open-source, MIT-licensed AI support agent you embed with one script tag. We're early: no wall of logos, no review-site score to lean on. The one trust asset available to a company like ours is honesty a stranger can verify, and claim drift burns it invisibly — usually midway through a technical evaluation, when a developer checks.
To be clear about how the drift accumulated: nobody sat down and decided to fabricate features. Copy got written against a roadmap, the code took a different route, and nobody ever diffs the homepage against the repo. Drift never feels like lying from the inside. The visitor reading the page can't tell the difference, so functionally it is.
Every fix below shipped as [one public pull request](https://github.com/theopenco/llmchat/pull/156). You can read each diff.
## The six claims that failed the diff
**1. "Run any model or provider."**
The model picker is a curated catalog of web-search-capable models, filtered from a generated snapshot of our gateway's catalog. Higher tiers unlock more of the list. "Any model" was aspirational copy for a picker that refuses to even boot with an empty list:
```ts
// packages/shared/src/models.ts
// Loud failure, never a blank picker: if the generated snapshot is ever empty
// (a botched regen), fail at import rather than silently offer no models.
if (WEB_SEARCH_MODELS.length === 0) {
throw new Error(
"WEB_SEARCH_MODELS is empty — run `pnpm gen:web-search-models` to regenerate from @llmgateway/models",
);
}
```
A curated catalog is a defensible design choice. We rewrote the copy to describe it, because it's what you get.
**2. "How many exchanges before the bot hands off."**
That's how the site described the escalation threshold: as if the agent decides, at some configured point, to hand the conversation over. The setting behind that copy is a per-project message threshold, and what it does is reveal a "Talk to a human" button after N messages. The visitor decides. A counter, not a judgment. The copy now describes the threshold.
**3. "Answers when it can. Hands off when it can't."**
Tidy copy, and it claims the AI monitors its own confidence and bails out the moment it's unsure. Great feature. We don't have it. What actually exists is visitor-initiated hand-off: the threshold button above, plus pattern detection for a visitor explicitly asking for a person. The detector's own doc comment is more honest than our homepage was:
```ts
// packages/widget/src/escalation-intent.ts
/**
* Detects a visitor explicitly asking for a human, so the "Talk to a human"
* CTA can surface immediately instead of waiting for the message-count
* threshold.
*
* Matching leans toward recall over precision: a match only REVEALS the
* escalate button (the visitor still has to click it), so a rare false
* positive costs one extra affordance while a false negative traps a
* frustrated visitor with the bot.
* …
*/
```
Even the explicit-ask path only reveals a button. The visitor clicks it. Nothing anywhere in the codebase asks the model how confident it feels. This one stung the most, because the fake version sounds smarter and we'd absorbed it into how we described the product out loud.
**4. Infrastructure attributed to the wrong vendor.**
Our comparison pages said "Fully self-hostable on Cloudflare infrastructure — D1, KV, and workerd." We don't run on Cloudflare. We run on serverless workerd via a platform called Ploy, and the same copy now reads "serverless workerd via Ploy (D1-compatible SQLite + KV state)." Same runtime family, wrong vendor. Nobody sues over this one, but a reader who catches the infrastructure paragraph being wrong has no reason to trust the security paragraph.
**5. "Self-host free, full feature set."**
The worst one, because it bent the promise open-source people actually check. Self-hosting is free forever with your own LLM keys; that part was true. What the copy skipped: a fresh self-hosted install resolves to a locked plan unless an environment allowlist is set. The unlock is a few lines of config parsing:
```ts
// apps/api/src/lib/plan.ts
export function internalEmails(env: Env): string[] {
const raw = env.vars.INTERNAL_ACCOUNT_EMAILS;
if (!raw) return [];
return raw
.split(",")
.map((s) => s.trim())
.filter(Boolean);
}
```
Unset, no workspace is exempt, and software running on your own server greets you with a paywall you had no way to anticipate. The immediate fix was documentation: [the docs](https://docs.clankersupport.com) now tell you the unlock exists and how to set it, instead of letting you discover the lock the hard way. Documenting an awkward mechanism beats hiding it, and it bought us time to decide what the mechanism should become.
**6. SSO/SAML and audit logs listed on the enterprise tier.**
Display copy with zero code behind it. We don't mean "beta" or "partially built" — grep the repo for SAML and the only hits are the copy itself. Both are now labeled as roadmap. Had an enterprise buyer asked for an audit-log demo, the demo would have been us typing very fast in another room.
A seventh came from an internal doc rather than the site: a per-plan member cap we claimed to enforce. The entitlement number exists in the billing config, but there's no invite endpoint, so there's nothing to gate. A limit with no enforcement path is a wish with a number on it. Claim dropped.
## One page undersold us
The audit cut the other way exactly once. Our own Chatwoot comparison listed "fully open-source (MIT license) — read, fork, and contribute" as the competitor's advantage, implying we weren't. We are, and the LICENSE file has been in the repo the whole time. That fix went into the same pull request as the six above.
Finding it reframed the whole exercise. The target is agreement between the site and the repo, and undersell is the same defect as oversell: the two disagree. Treat both directions as bugs and the audit stops feeling like penance and starts feeling like ordinary QA, which is what got it finished.
## The re-pass found 18 more, four written during the fix
After the first fix pass we felt pretty good about ourselves. Then we ran an adversarial review: a two-agent panel, LLMs given the repo, whose only job was to attack the corrected copy against the codebase.
They found 14 missed instances of the same six overclaims, spread across pages, docs, and meta descriptions the first pass never opened.
Worse, they found 4 brand-new overclaims in the replacement copy itself. Text written during a truth-telling exercise, by people actively trying to be accurate, still drifted optimistic within the same edit session. We'd swap a false claim for a true one and unconsciously round it up while typing.
Marketing language has gravity: every sentence wants to be slightly more impressive than the facts, and you don't feel the pull while writing. The only countermeasure we've found is a second pass by someone, or something, that doesn't share your incentives.
## Diff your homepage against your repo
The repeatable version fits in an afternoon, and you don't have to read code to run it. Each step is an ask you can hand to your team as one sentence.
1. **Extract every falsifiable claim.** Scrape your homepage, pricing page, feature pages, and docs. Pull out every sentence that asserts something checkable: a feature exists, a limit is enforced, a platform is used, a license applies. Ignore vibes ("delightful"), keep facts ("supports SSO").
2. **Assign each claim a code location.** For every claim, ask: which file or endpoint makes this true? "Any model" should point at the model list. "Audit logs" should point at an audit-log table. A claim nobody can point anywhere is a finding, and whoever wrote the feature can answer the question in about a minute.
3. **Classify: true / roadmap / false / understated.** A roadmap item is fine as long as it's labeled as one. Anything shipped-but-unclaimed goes in the understated bucket, and it gets fixed in the same pass.
4. **Fix everything in one public pull request.** The public part is the point. A private cleanup earns nothing; a public one is evidence you can hand a skeptical visitor for years.
5. **Re-audit the fix adversarially.** This is the step everyone skips, and we nearly did. It produced 18 of our findings — four of which we authored during the fix itself. A colleague, an advisor, an LLM given the repo and told to attack: anyone whose job is to disagree with your copy.
Then put it on a calendar, because drift regrows and every new landing page restarts the clock. The cost of skipping the whole exercise never appears as a line item. It arrives as an evaluation that quietly churned when a developer caught claim number four, or an enterprise call where someone asks to see the audit log.
We've run a related exercise on our search presence before, in [our AI SEO audit checklist](https://clankersupport.com/blog/ai-seo-audit-checklist). Don't blur the two: that one audits how your pages present claims to search and answer engines; this one audits the product claims themselves against code.
## The cheapest credibility available
The product carries the same rule at a smaller scale. Where the dashboard has no real value for a metric, it renders an em dash rather than a guessed number, and the inbox stats' doc comment spells the policy out:
```tsx
// apps/dashboard/src/app/inbox/_components/InboxStats.tsx
/**
* …The avg rating is the mean CSAT across rated conversations only, shown
* as "—" when none are rated (never NaN). While the aggregate is loading,
* values render as "—".
*/
```
The usage meter in the sidebar goes further and renders nothing at all until the number resolves — its comment reads "(no fabricated zero)". That's the same rule our homepage broke six ways, applied at the level of a single stat card.
What a company at our stage can offer is verifiability: an MIT repo you can read, a public pull request where we corrected our own marketing, docs that explain the awkward parts like the self-host unlock instead of burying them. A skeptical developer can check every word of this post against the diffs in about ten minutes, and that checkability is worth more to us than any adjective we could have kept.
If you're early and your homepage promises things your repo can't cash, you're spending trust you haven't minted yet. Diff the site against the repo and fix in both directions, where people can watch. It's an afternoon of mildly humiliating work, and it's the cheapest credibility you'll ever buy.
## FAQ
### How do you audit marketing claims against a codebase?
Extract every falsifiable claim from your homepage, pricing page, and docs; assign each one the file or endpoint that would make it true; classify each as true, roadmap, false, or understated; fix everything in one public pull request; then have someone adversarial re-audit the fix. The last step matters most — our re-pass caught overclaims written during the fix itself.
### What is claim drift?
Claim drift is the gap that grows between what a marketing site says and what the codebase does. It rarely starts as a lie: copy gets written against a roadmap, the product takes a different route, and nobody re-checks the copy. The visitor can't distinguish drift from fabrication, so it costs the same trust.
### Should the fixes be public?
If the product is open source, yes. A private cleanup earns nothing, while a public pull request is standing evidence you can hand a skeptical evaluator years later. It also raises the cost of future drift, which is half the point.
# Two messages, one sequence number: the concurrency bug that never threw
URL: https://clankersupport.com/blog/two-messages-same-sequence-number
Published: 2026-07-26
Category: Engineering
An escalation marker and the AI reply it interrupted landed in the same conversation with the same sequence number. No exception, no log line, both inserts succeeded, and the thread quietly rendered out of order. This is the read-then-write bug that passes every test, the three-part fix whose deploy order is the point, and the five-second code search that tells you whether your codebase has it too.
One day, two messages in the same support conversation both had `sequence = 2`. Both inserts succeeded and nothing was logged; everything downstream simply picked an order. The colliding pair says it all: the "Visitor requested a human operator" marker and the AI answer the visitor was escalating away from, and depending on which order a client picked, the request for a human rendered before or after the answer that prompted it. That `sequence` integer is load-bearing in Clanker Support: the widget orders the thread by it, the dashboard inbox orders by it, unread tracking uses it as a high-water mark. For a support product, that is close to the worst available failure. The conversation thread is the artifact your customer trusts, and a garbled support thread fails the way a garbled bank statement does: the customer stops believing the record.
The expensive property of this bug class is who finds it. It never appears in tests, never appears in logs, and the corruption compounds quietly, so the first person positioned to notice is a customer reading a thread that makes no sense — long after the writes that caused it. We caught it in our own testing before it cost anyone anything; we're early, and that's the cheapest possible place to catch it. Nothing about the race changes at scale except how much data it quietly ruins first. The product is open source ([theopenco/llmchat](https://github.com/theopenco/llmchat)), so everything below links to real diffs.
## Two sins, one write path
Every writer in the system did some version of this: read the conversation, do work, then write a message whose sequence was computed from the earlier read. Simplified:
```ts
// BEFORE: read early, write late
const convo = await db.query.conversation.findFirst({
where: eq(conversation.id, conversationId),
});
// ... 5–20 seconds of LLM streaming happens here ...
await db.insert(message).values({
conversationId,
content,
sequence: convo.messageCount + 1, // computed from a read made earlier
});
await db.update(conversation).set({ messageCount: convo.messageCount + 1 }); // absolute assignment
```
Sin one: `sequence` is computed from `messageCount` as it stood when this request read it. If any other writer inserts between the read and the insert, both writers saw `messageCount = 3` and both write `sequence = 4`. SQLite takes both rows without complaint.
Sin two: the count update is an absolute assignment. Writer A sets `messageCount = 4`; writer B, working from the same stale read, also sets it to 4. Two messages arrived and the count moved by one. The damage compounds from there, because the next writer derives its sequence from a count that's already wrong, so one race seeds the next.
We had no shortage of concurrent writers, either. The chat handler persists the assistant's reply after the stream finishes, inside `waitUntil`, seconds after it read the conversation. An operator can reply from the dashboard inbox at any moment. Internal notes land. An escalation writes a system message. An inbound email reply arrives through a webhook. Five independent writers, all doing read-then-write against the same counter.
## Why every test passed
Our tests exercised each writer in isolation: insert a user message, assert sequence 1; insert a reply, assert sequence 2. Green across the board.
The race needs two writers interleaved inside the same window, and the widest window in the whole system is the one no unit test reproduces: the assistant-persist that runs after an LLM stream completes. In tests the "stream" resolves instantly with nothing else running. In production it takes 5 to 20 seconds, and the true trigger is mundane: the visitor hits "Talk to a human" mid-stream (the exact collision we found), or an operator replies from the inbox while the agent is still streaming. Either one lands a write inside the gap every single time it happens.
So tests pass because they're sequential, and production fails because it isn't. No quantity of extra tests fixes that cleanly. The gap itself has to go.
## The fix shipped in three steps, and the order is the point
You can't just add a unique index: the old writers are still running while you deploy, and the index build fails outright if duplicates exist. You can't just fix the writers either: existing duplicate rows stay corrupted, and any future regression goes back to being silent. So the fix went out as three deploys, strictly ordered.
### Clean the data first
The backfill ([the public pull request](https://github.com/theopenco/llmchat/pull/161)) finds every conversation with at least one duplicated `(conversation_id, sequence)` pair, renumbers all of that conversation's messages with `ROW_NUMBER()`, then trues up the drifted `message_count`. Lightly trimmed:
```sql
CREATE TABLE _seq_backfill AS
SELECT id AS mid,
ROW_NUMBER() OVER (PARTITION BY conversation_id
ORDER BY sequence, created_at, id) AS new_seq
FROM message
WHERE conversation_id IN (
SELECT conversation_id FROM message
GROUP BY conversation_id, sequence HAVING COUNT(*) > 1);
UPDATE message
SET sequence = (SELECT new_seq FROM _seq_backfill WHERE mid = message.id)
WHERE id IN (SELECT mid FROM _seq_backfill);
DROP TABLE _seq_backfill;
UPDATE conversation
SET message_count = (SELECT COUNT(*) FROM message m WHERE m.conversation_id = conversation.id)
WHERE message_count <> (SELECT COUNT(*) FROM message m WHERE m.conversation_id = conversation.id);
```
Two details worth stealing. The `ORDER BY sequence, created_at, id` makes the renumbering deterministic: ties on the duplicated sequence break by creation time, and ties on creation time break by id. That last tiebreaker matters more than it looks — our timestamps are unix seconds, so same-second writes are common, which is the whole bug. Run the backfill twice and you get the same answer. And it stages through a temp table rather than a correlated self-update, because renumbering a partition while you're reading it is how you end up with a backfill you can't reason about.
### One writer, and the database allocates the key
[The writer refactor](https://github.com/theopenco/llmchat/pull/162) replaces every inline insert-and-bump in every route with one function: `insertMessage()` in `apps/api/src/lib/messages.ts`. Chat persist, operator reply, notes, escalation markers, inbound email replies: all of them now go through it.
The core move is that the sequence is allocated by a scalar subquery inside the INSERT itself:
```ts
db(env)
.insert(message)
.values({
conversationId: input.conversationId,
role: input.role,
content: input.content,
// ...
sequence: sql`(SELECT COALESCE(MAX(${message.sequence}), 0) + 1 FROM ${message} WHERE ${message.conversationId} = ${input.conversationId})`,
})
.returning();
```
That's the entire trick. SQLite (and D1, which is SQLite at the edge) serialize writers per statement, so `MAX(sequence) + 1` and the insert are atomic. There is no gap between reading the current max and writing the next value, because they're the same statement. The database allocates the ordering key, and the application never holds it in a variable where it can go stale.
The counter fix rides along:
```ts
const bumped = await db(env)
.update(conversation)
.set({
messageCount: sql`${conversation.messageCount} + 1`,
updatedAt: new Date(),
})
.where(eq(conversation.id, input.conversationId))
.returning({ messageCount: conversation.messageCount });
```
`message_count = message_count + 1` is commutative: two concurrent bumps produce +2 no matter how they interleave, where two absolute assignments produced +1. The `RETURNING` clause hands the post-bump count back to callers so nobody is tempted to re-derive it from a sequence number.
### The unique index that makes any regression loud
The last deploy adds [the unique index](https://github.com/theopenco/llmchat/pull/163) that makes any regression loud:
```sql
CREATE UNIQUE INDEX IF NOT EXISTS message_conv_seq_uidx
ON message (conversation_id, sequence);
```
And the part we'd get wrong if we did this again without notes: that migration re-runs the exact same dedupe backfill, immediately before creating the index, in the same migration. Between the backfill deploy and the writer deploy, the old racy writers were still live in production. Any duplicate they minted in that window would make the index build fail and take the whole deploy down with it. The re-dedupe costs nothing when the data is already clean and saves the deploy when it isn't. (We're rigid about deploy ordering around risky schema changes in general — it's the same discipline that kept [an "always safe" additive column from breaking every login](https://clankersupport.com/blog/additive-migration-almost-broke-every-login).)
Once the new writer is live, the index should never fire. It exists so that if someone adds an inline insert with a precomputed sequence eight months from now, the result is a constraint error in the logs instead of six more weeks of silently shuffled threads. Downgrading a bug from silent corruption to loud error is most of the value of the whole exercise.
## Drizzle buries the error your retry path needs on `.cause`
`insertMessage()` has a retry path: if an insert ever trips the unique index, it retries once, and the re-run subquery naturally picks the next free slot. To do that it has to detect a unique-constraint violation, and our first attempt was the obvious one:
```ts
// looks right, never matches in prod
if (err instanceof Error && /unique constraint failed/i.test(err.message)) {
```
It never matches. Drizzle 0.45 wraps every driver error in a `DrizzleQueryError` whose message reads `"Failed query: insert into message ..."`. The actual `UNIQUE constraint failed: message.conversation_id, message.sequence` text lives on `err.cause`, one level down. A message-only check compiles, passes any test that fakes the error, and silently classifies every real violation as an unknown error to rethrow. The retry path becomes dead code, and you find out the day the tripwire fires and nothing retries.
The version that works walks the cause chain, with a cycle guard because `cause` can technically point anywhere:
```ts
function isUniqueViolation(err: unknown): boolean {
const seen = new Set();
for (let e: unknown = err; e && !seen.has(e); ) {
seen.add(e);
const msg = e instanceof Error ? e.message : String(e);
if (/unique constraint failed/i.test(msg)) {
return true;
}
e = (e as { cause?: unknown }).cause;
}
return false;
}
```
If you classify database errors through any ORM, check whether the string you match lives on `.message` or on `.cause.message`. `Error.cause` chains have been standard since ES2022, and most error-classification code we've read predates them.
## The five-second audit
The rule underneath all of this: never derive an ordering key from state you read earlier in the request. Not a sequence number, and not a version counter either. The moment the value leaves the database and sits in a variable it's a snapshot, and every millisecond between read and write is a window some other writer will eventually hit. Streams, webhooks, and background jobs stretch those windows to seconds.
You don't need to write code to check your own product for this today. Ask whoever owns the backend to search the codebase for `count + 1` or `position + 1` computed in application code and written back later. The search takes about five seconds, and each hit is this bug wearing different clothes. Then ask two follow-ups: is there one writer function for that key, and is there a database constraint that would make a violation loud? If either answer is no, you have this bug on a timer.
The repair, when you need it, is the same recipe in the same order:
1. **Backfill first**, deterministically and idempotently, so the data is clean.
2. **Let the database allocate the key atomically**, in the same statement as the write, through one writer function with no exceptions.
3. **Add the constraint last**, re-cleaning immediately before you build it, so the bug class becomes impossible rather than unlikely, and any regression is a loud error instead of quiet drift.
All three diffs are small enough to read in one sitting, and every one of them is public on [our repo](https://github.com/theopenco/llmchat).
## FAQ
### How do you safely generate a per-group sequence number in SQLite or D1?
Allocate it inside the INSERT itself with a scalar subquery: `sequence = (SELECT COALESCE(MAX(sequence), 0) + 1 FROM message WHERE conversation_id = ?)`. SQLite serializes writers per statement, so reading the current max and writing the next value happen atomically. Any pattern that reads a counter into application code first has a race window, and streaming or background work stretches that window to seconds.
### Why isn't a unique index alone enough to fix a sequence race?
Deploys aren't instantaneous. While the index migration runs, older application code with the racy writer is still serving traffic, and any existing duplicates make the index build fail outright. Clean the data first, route every write through one atomic writer, then create the index, re-running the dedupe immediately before it to absorb duplicates minted between deploys.
### Why doesn't err.message contain "UNIQUE constraint failed" with Drizzle?
Drizzle 0.45 wraps driver errors in a `DrizzleQueryError` whose message describes the failed query; the constraint text lives on `err.cause`. Walk the cause chain (with a cycle guard) when classifying database errors, or your retry and fallback paths will silently never run.
# Every hosted plan now starts with a 7-day free trial
URL: https://clankersupport.com/blog/7-day-free-trial
Published: 2026-07-20
Category: Announcements
New hosted Clanker Support subscriptions now begin with 7 free days — the full plan, every feature, applied automatically at checkout. Nothing is charged until the trial ends, you can cancel anytime, and a 14-day money-back guarantee backs it all up.
As of last week, every new hosted Clanker Support subscription starts with a 7-day free trial. It applies automatically at checkout — no promo code, nothing to hunt for. Pick a plan, and for seven days you use it for free.
Here's the honest version of why. Our hosted product has never had a free tier — that's by design, because self-hosting is the free version and always will be. But it meant the first thing a new customer met was a paywall, before the product had said a word for itself. We didn't like that first impression, so we replaced it: now the first thing you meet is the full product, free for a week.
## The trial is the whole plan, not a preview
The seven days are the plan you picked, from day one. Every feature, the full response allowance — 2,000 AI responses on Starter, 12,000 on Growth, 50,000 on Scale. There's no demo mode, no locked features, no "upgrade to unlock" halfway through.
That matters because of what a week is actually enough for. Install is one script tag, and most teams are live in about five minutes. Point the agent at your docs and knowledge sources, and it starts answering your real customers — escalating to your team instead of guessing when it can't. By day seven you're not evaluating our product anymore. You're deciding whether to switch off something that's already handling your support.
## Why we ask for a card
We do ask for a card at checkout, and we'd rather explain that than hide it.
Nothing is charged during the trial. Usage during the trial is never billed — not on day seven, not ever. The first charge happens only when the trial ends and you've chosen to stay.
The card is there so the week ends cleanly either way. If Clanker Support has earned its place, your subscription continues without you re-entering anything or losing a day of coverage. If it hasn't, cancelling is self-serve from your billing settings — through the Stripe billing portal, before day seven, and you pay nothing.
## When the seven days end
If you stay, your plan simply continues and your first charge goes through. And you're still covered: every hosted plan comes with a 14-day money-back guarantee, you can cancel anytime, and there are no contracts.
Stack that up and the arrangement is deliberately lopsided. Seven free days, no charge until the trial ends, self-serve cancellation from billing settings, money-back guarantee after that. We carry the risk of you trying Clanker Support. You don't.
## Already a customer?
One thing we want to be straightforward about: switching between paid tiers doesn't restart a trial. The trial is for workspaces starting their first subscription — if you're already with us and move from Starter to Growth, the change applies right away, without a second free week.
We think that's the fair version. The trial exists so new customers can see the product work before paying, not as a loop to be replayed. Existing customers already know what they're paying for.
## Self-hosting stays free
Nothing about this changes the open-source side. Clanker Support is open source, and self-hosting is free forever — bring your own LLM keys and run the whole thing on your own infrastructure. The hosted plans are for teams who'd rather we operate it; the trial just means trying that now costs nothing either.
## How to start
1. **See it working first, with no signup.** The [live demo](https://showcase.clankersupport.com) is the real widget running in your browser — open it and ask it something.
2. **Pick a plan on [/pricing](https://clankersupport.com/pricing).** We'd suggest starting on Starter — $19 a month, 2,000 AI responses, no per-seat fees — and switching tiers later if you outgrow it. Annual billing gets you two months free. Your 7-day trial starts at checkout, automatically.
3. **Drop in the script tag.** One script tag on your site; most teams are live in about five minutes. Then let it take your real conversations for a week.
## Questions you might have
### Why do you require a card for a free trial?
So the trial ends cleanly. Nothing is charged during the seven days — if you stay, your plan continues without interruption or re-entering details; if you don't, you cancel yourself from billing settings and pay nothing.
### What happens when the 7 days end?
Your subscription begins and your first charge goes through. You're still protected by the 14-day money-back guarantee, and you can cancel anytime — there are no contracts.
### Does usage during the trial cost anything?
No. Trial usage is never billed, no matter how much of your plan's allowance you use.
### I already subscribe — do I get a trial if I switch plans?
No — tier changes apply immediately without a new trial. The trial is for workspaces starting their first hosted subscription.
### Is self-hosting still free?
Yes, forever. Open source, your own infrastructure, your own LLM keys.
## Start wherever you're comfortable
There's a ladder here, and you can stop on any rung. Try the [live demo](https://showcase.clankersupport.com) with no signup at all. When you're curious what it does with your docs and your customers, take the 7 free days. And if by day seven your support agent already feels like yours — that's the point at which staying is the easy decision.
Pick your plan at [/pricing](https://clankersupport.com/pricing). The trial starts the moment you do.
# 'Additive columns are always safe' is wrong on Drizzle, Prisma, and preview deploys
URL: https://clankersupport.com/blog/additive-migration-almost-broke-every-login
Published: 2026-07-20
Category: Engineering
Everyone agrees an additive column is the one schema change that can't hurt you. Then a one-line ALTER TABLE on our user table turned out to be capable of 500ing every authenticated request — because Drizzle projects every mapped column, Better Auth reads the user table on every session check, and preview deploys skip migrations. Here is the two-PR discipline we ship our riskiest schema changes with, and the one column we keep out of the ORM entirely.
Migration `0017_user_role.sql` in our repo is one line of SQL under eighteen lines of comment, and the comment is the interesting part: it explains why the column it adds must never appear in our ORM schema. Declared the way every tutorial shows, that one additive column would have 500'd every authenticated request on any database that hadn't run the migration yet. "Additive nullable columns are always safe" is received wisdom, and on a modern stack it is false: ORMs like Drizzle and Prisma enumerate every mapped column on every SELECT, preview deploys run new code against old schemas, and auth libraries query your user table on every request — so one unmigrated column can fail every query that touches its table.
To be precise about what actually happened, because this is easy to overclaim: no production outage. What we had was a string of preview deploys 500ing on columns that existed only in code, and one near-miss — that `role` column — where the same mechanism pointed straight at the auth hot path. This is the failure mode we kept almost shipping, and the discipline that stopped it. Clanker Support is open source ([theopenco/llmchat](https://github.com/theopenco/llmchat)), so every file, commit, and PR below is public.
## Why "additive columns are always safe" became folklore
The belief was earned, and it predates ORMs. Adding a nullable column — or a `NOT NULL` column with a default, the other blessed shape — rewrites nothing. In SQLite it's a metadata change; Postgres has done the same for defaulted columns since version 11. No lock, no backfill, no data risk. And the load-bearing clause: old code ignores columns it doesn't know about. `SELECT id, email FROM user` does not care what else the table grew this week.
That clause was true when column lists were handwritten. Then they started being generated from schema files that ship with the application code, and the clause quietly inverted: now the code can know about a column before the database does. Additive migrations are safe when the schema definition trails the database. They are dangerous when it leads — and on a modern deploy pipeline, it leads all the time.
## Your ORM selects every column you map
Drizzle does not emit `SELECT *`. An unprojected `select().from(user)` expands to an explicit list of every column mapped on the table object in `packages/db/src/schema.ts` — for our `user` table, all seven of them, by name. Map an eighth column that the database doesn't have yet and every one of those queries throws `no such column`, including queries whose calling code never reads the new field. The migration isn't what breaks. Reads that predate the feature are what break.
This is not a Drizzle quirk. Prisma's generated client selects every scalar field by default unless you pass `select`. Rails people know the mirror image of this rule from column _removal_ — you set `ignored_columns` before dropping, because the schema cache still names the column — but explicit-projection ORMs make column _addition_ just as directional. Any ORM that generates its column lists from a checked-in schema has the same property: the table's every reader is coupled to the schema file's most recent line.
## Preview deploys run new code against old schemas
For the mismatch to bite, some environment has to serve new code against an old database. Our platform hands us that environment on every branch: production deploys apply the migrations in `apps/api/migrations/`, preview deploys don't. A branch that adds a column and maps it in `schema.ts` gets a preview whose code names a column its database will never have. We watched exactly this — previews returning 500s for a column only prod would ever get — which is the cheap version of the lesson, paid in red preview checks instead of pages.
Previews are the guaranteed case, not the only one. A Vercel preview pointed at a shared staging database has the same gap, and so does a Neon branch-per-preview setup where the branch was snapshotted before your migration existed. So does every self-hosted install that pulls your code before running your migrations — and, potentially, production itself during the deploy window: our platform's docs don't specify whether migrations apply before the new worker starts serving traffic, so we defend against both orders rather than betting on one. The environments differ; the shape is identical: the schema file leads, the database trails, and the ORM faults on the gap.
## The auth library that reads your user table on every request
Here is the multiplier that turns a red preview into a near-catastrophe. We were adding a platform-admin `role` to `user` for our internal admin console — a column that gates three admin routes only our own team calls. Its natural blast radius is approximately zero. Its actual blast radius, had we mapped it in Drizzle, is documented in a comment we now keep on the table itself:
```ts
// packages/db/src/schema.ts
// NOTE: the PLATFORM-admin role column (migration 0017_user_role.sql) is
// deliberately NOT modeled on this Drizzle table. Better Auth's Drizzle
// adapter loads the session user with an UNPROJECTED `select().from(user)`
// (every column of this table object) on every getSession, so declaring
// `role` here would make that auth hot-path query reference a column a preview
// DB — which skips migrations — does not have, 500-ing ALL authenticated
// requests.
```
Better Auth's adapter hydrates the full user object on every session check. That is not a bug, and we want to be fair about it: the adapter can't know which subset of columns your application needs, so returning the whole row is a reasonable contract, and Better Auth core has otherwise been solid for us. But it means the `user` table's column list is load-bearing for 100% of authenticated traffic, and any column you add to the mapping joins the hottest path in the system the moment you commit it. A column nobody reads would have taken down the inbox, the settings pages, sign-in — everything behind a session — on any lagging database.
The transferable lesson is not "audit your auth library". It's that your dependencies' query shapes are production behavior you own. You can read every line of your own code and still not know which of your tables gets an unprojected read per request.
## Two PRs per schema change: migrate before you serve
The fix is old — the expand/contract pattern, in miniature. Every schema change on the tables our hottest paths read — `message`, `conversation`, `user` — ships as two pull requests. Phase 1 is the `ALTER TABLE`, its comment, and a seed-contract test; `schema.ts` is deliberately untouched, so no deployable code can name the column under any deploy ordering, in any environment. Phase 2 — the Drizzle mapping, the endpoints, the UI — merges only after phase 1 is live in production. We should admit the discipline is risk-scoped, not universal: lower-stakes columns on `project` — settings fields read by a handful of routes — still went out as single PRs (`0016`, `0020`, `0021`), which is a bet that nobody needs that table's preview to work that week. The hot tables don't get the bet. The migration files carry the reasoning in full, and they've become the best documentation in the repo:
```sql
-- apps/api/migrations/0022_message_reply_to.sql
-- PHASE 1 of 2 (deliberate migrate-before-serve split, mirroring 0014/0015).
-- Ploy's deploy ordering between "apply migrations" and "new Worker serves
-- traffic" is undocumented, so this PR ships ONLY the column — schema.ts is
-- intentionally NOT changed, so the live Worker never SELECTs a column that might
-- not exist yet (drizzle projects every column of `message`, and /v1/chat +
-- /v1/messages read it on the hottest paths; no read can 500 under any ordering).
ALTER TABLE `message` ADD COLUMN `reply_to_message_id` text;
```
We've run the split three times so far:
- **`0014` conversation summaries** — migration in PR #76, feature in PR #77, merged 24 minutes apart on June 22.
- **`0015` resolve attribution** — migration in PR #91, feature in PR #92, 41 minutes apart on June 29.
- **`0022` quote-reply** — migration in PR #142 on July 12, feature in PR #143 two hours later that night, once the column was confirmed live (the phase-2 commit message records that "migration 0022 is already live in prod").
That's the honest cost accounting: two PRs instead of one, and between 24 minutes and a couple of hours of waiting. Against that, for the changes we split, there is no deployable commit where code names the column before its migration is live in prod — the class of 500 becomes unrepresentable. The long comments are part of the discipline, not decoration: a one-line `ALTER TABLE` in its own PR looks like pointless ceremony six months later, and the comment is what stops the next person from helpfully collapsing it back into one.
## The column we keep out of the ORM entirely
`user.role` is the extreme case: phase 2 never came, on purpose. Because the `user` table's mapping is read unprojected on every request, the column stays out of `schema.ts` indefinitely, and the admin gate reads it with a raw SQL projection wrapped in a fallback (the `/admin/users` listing uses the same guarded shape):
```ts
// apps/api/src/middleware/admin.ts
try {
const rows = await db(c.env)
.select({ role: sql`role` })
.from(user)
.where(eq(user.id, userId))
.limit(1);
role = rows[0]?.role ?? null;
} catch {
role = null;
}
```
On a database without the `0017` migration, the read throws, `role` degrades to `null`, and the requester is a non-admin — a 403, never a 500. Failing toward least privilege is the correct direction for an admin gate anyway, so the defensive shape costs nothing.
One loose end remained: Drizzle's `query.user.findFirst` without a `columns` option is also a select-everything of the row. Commit `e287d58` swept the last two of those — billing checkout and project creation, each of which only ever read `.email` — down to explicit projections:
```ts
// apps/api/src/routes/billing.ts
const owner = await db(c.env).query.user.findFirst({
where: (u, { eq: e }) => e(u.id, userId),
columns: { email: true },
});
```
The invariant that fell out is easy to state and easy to check in review: exactly two queries in the codebase name `role` — the admin gate and the `/admin/users` listing — and both are wrapped in a try/catch that degrades to least privilege. Everything else touching `user` says which columns it wants. Once production and every preview convention has settled, the column can be folded into the schema like any other — the comment on the table says as much — but there is no hurry, because the current shape cannot break.
## When a schema change needs the two-PR split
| Change | Blast radius if code ships first | What to do |
| --------------------------------------------------------------------- | ----------------------------------------------------------- | --------------------------------------------------------------------------- |
| New table | None — no deployed code queries it | One PR is fine |
| New column, mapped in the ORM | Every ORM read of that table, via the generated column list | Two PRs: migration alone, then mapping + feature |
| New column on a table a dependency reads unprojected (auth, sessions) | Every authenticated request | Two PRs — or keep the column out of the ORM behind a guarded raw projection |
| Dropping a column | Deployed code still projects it during the rollout window | Two PRs in reverse: remove the mapping first, drop later |
| Self-hosted installs exist | You never control when they migrate | Treat every schema change as two-phase, always |
The first question to ask about any table is the one we didn't know to ask: who reads it that you didn't write? For `message` it was our own hot paths, which we could see. For `user` it was our auth library, which we couldn't — until we looked at the queries it actually emits.
This is the second time that habit has paid for itself. When we [moved the backend to workerd](https://clankersupport.com/blog/cloudflare-workers-every-node-sdk-broke), the lesson was to audit transitive dependencies, not imports — the package that broke your deploy wasn't the one you installed. This one is the same lesson at the database layer: audit your dependencies' query shapes, not just your own reads. The code you didn't write is still your production behavior. The migration comments in `apps/api/migrations/` are all public if you want the long-form version; they're better reading than most of our docs.
## FAQ
### Is adding a nullable column a safe migration?
Only if no deployed code selects it before it exists. The database operation is safe; the hazard is your ORM. Drizzle and Prisma generate explicit column lists from the schema definition, so a mapped-but-unmigrated column fails every read of that table — in previews that skip migrations, in self-hosted installs, and during the deploy window itself.
### Why does my preview deploy fail with "no such column"?
Your branch maps a new column in the ORM schema, but the preview database never ran the branch's migration. The ORM names the column on every SELECT of that table, and the database rejects it. Ship the migration in its own PR first, and add the ORM mapping only after the column is live everywhere that serves traffic.
### What is the expand/contract migration pattern?
Splitting a schema change so that every deployed version of the code works against both the old and new schema: add the column first (expand), deploy, then ship the code that uses it; for removals, delete the code references first, then drop the column (contract). Our two-PR split is the smallest useful version of it.
# A 9-point AI SEO audit checklist, run on our own site first
URL: https://clankersupport.com/blog/ai-seo-audit-checklist
Published: 2026-07-20
Author: Ismail Ghallou
Category: Guides
Next.js merges route metadata shallowly, so every page built with our metadata helper — unless it passed a cover image of its own — shipped a Twitter card with no image for about a month. That was finding one of the SEO and AI-SEO audit we ran across our five domains this month, in two passes. Here is the whole audit as a checklist you can run on your own site, including the four findings that surprised us.
For about a month, every page on our marketing site that used our own metadata helper without an image of its own — `/pricing`, every `/vs/*` comparison, every `/features/*` page — shipped a `summary_large_image` Twitter card with no image. Not a broken image. No image at all, on pages whose metadata we thought a shared helper had made uniform. The cause is a Next.js behavior that hits any site combining a root Open Graph image with page-level metadata — which is most maturing Next.js sites — and we'll get to it first because it earned its place at the top of the checklist.
We found it while auditing our five web surfaces (marketing, docs, dashboard, showcase, admin). An AI-SEO audit checks the machine-facing surface of your site twice: once for search crawlers — robots rules, canonicals, sitemaps, Open Graph, structured data, index hygiene — and once for AI answer engines — crawler access for GPTBot, ClaudeBot and PerplexityBot, `llms.txt` and `llms-full.txt` files, and self-contained extractable answers. The output is a pass/fail list per domain you operate, including the domains your product mints pages on.
The audit landed in two waves: a first pass of fixes straight to main on July 6–8, then the cross-domain wave that merged as [PR #148](https://github.com/theopenco/llmchat/pull/148) on July 17 — that one touched marketing, docs, showcase, admin and the api. Everything is public in [the repo](https://github.com/theopenco/llmchat), so every finding links to real code. Here is the checklist, then the four findings that were genuinely non-obvious.
## The checklist
Run each question against every domain you operate — including subdomains you forgot you had.
| # | Check | Ask yourself | Us, before the audit |
| --- | -------------------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| 1 | Social card images | Does the **rendered HTML** of every page contain `og:image` — not just the homepage? | Fail — every helper-built page without its own cover had none |
| 2 | noindex reachability | Can crawlers fetch your noindex'd pages? A robots Disallow hides the tag | Fail — the dashboard got noindex + robots in wave one; admin's robots.txt came in wave two |
| 3 | Product-minted URLs | Does every per-customer URL your product serves send noindex? | Fail — `/embed/:key` was indexable |
| 4 | AI crawler access | Do any robots rules block GPTBot, ClaudeBot, PerplexityBot, or Google-Extended? | Pass — nothing blocked (now deliberate) |
| 5 | llms.txt files | Do you publish a link map (`/llms.txt`) and a full-content file (`/llms-full.txt`)? | Half — link map yes, full-content no |
| 6 | Secondary domains | Does every subdomain have robots, a sitemap, canonicals, OG tags, and structured data? | Fail — the docs subdomain had none of the five |
| 7 | Title lengths | Do your titles survive the ~60-character SERP cutoff? | Fail — a brand suffix pushed every post past it |
| 8 | Sitemap honesty | Is `lastModified` a real content date or a build timestamp? | Fail — build timestamp on every entry |
| 9 | Demo properties | Do showcase/demo sites have a `metadataBase`, a canonical, and a robots.txt that isn't a 404? | Fail — our showcase's robots.txt returned 404 |
Five domains, nine checks, and the only things that came through clean were the marketing site's JSON-LD, canonicals and llms.txt — the parts built deliberately. Even its sitemap was lying about dates, and the RSS feed it should have advertised didn't exist until the first wave added it. Everything that grew organically had a gap.
## Why your Next.js pages have no og:image
The mechanism: Next.js merges route metadata **shallowly**. When a page exports its own `openGraph` object, it doesn't extend the layout's — it replaces it wholesale, and the same applies to `twitter` and `alternates`. Our root layout ships a site-wide OG cover via the `opengraph-image.png` file convention, and we assumed that cover reached every page. It reached exactly the pages that declared no `openGraph` of their own. Every page that called our `pageMeta` helper without an explicit image — `/pricing`, the comparisons, the feature pages, the tools — wiped the image out in the same stroke, and these were the pages we had put the most metadata care into. Blog posts survived only because they pass their own cover image to the same helper.
The failure is invisible in the browser and in most SEO tooling, because the pages still had titles, descriptions and canonicals. What they emitted was a `twitter:card` of `summary_large_image` with no image behind it, so every share of `/pricing` or a comparison page rendered as a bare text stub. We only caught it by grepping the prerendered HTML for `og:image` during the audit.
The fix is a constant and a default: put the image inside the helper, so no caller can forget it. The comment in `apps/marketing/src/lib/seo.ts` is the whole postmortem:
```ts
// apps/marketing/src/lib/seo.ts
/** The site-wide OG cover (the app/opengraph-image.png file convention route).
* Explicit fallback because a page-level `openGraph` object replaces the
* layout's resolved metadata wholesale (Next merges shallowly) — without this,
* every pageMeta page shipped no og:image/twitter:image at all. */
const DEFAULT_OG_IMAGE = "/opengraph-image.png";
```
The same shallow merge had already bitten us once in the same file, back in the first wave — the `alternates` block re-declares the RSS `` on every page, because a page-level canonical would otherwise delete the feed reference the layout set. Same behavior, different casualty.
Check number 1, generalized: if you use any metadata helper or page-level `openGraph` in a Next.js App Router site with a root OG image, view source on a non-homepage page and search for `og:image`. This combination — root image plus page-level metadata — is the default shape of a maturing Next.js site, which is why we're comfortable saying the bug is widespread. It costs you every social and chat-app share silently, and no console warns you.
## robots.txt Disallow doesn't deindex — it does the opposite
The counterintuitive one. Our operator dashboard should never appear in search results, and the reflex is to write `Disallow: /` in its robots.txt. That reflex is wrong, and it's wrong in a way that leaves the pages _in_ the index.
A robots Disallow controls **crawling**, not **indexing**. A `noindex` meta tag controls indexing — but Google can only read the tag on pages it's allowed to fetch. Disallow a URL that anyone links to externally, and Google indexes it anyway, as a bare URL with no snippet ("Indexed, though blocked by robots.txt" in Search Console). Our marketing site links to the dashboard sign-in page, so a Disallow would have pinned that URL in the index with no crawlable signal to ever remove it.
So the dashboard does the opposite. `apps/dashboard/src/app/robots.ts`, rationale included:
```ts
// apps/dashboard/src/app/robots.ts
// Deliberately allow crawling: the layout serves a noindex robots meta on every
// page, and Google can only see that tag on pages it's allowed to fetch. A
// Disallow here would leave externally-linked URLs (the marketing site links to
// sign-in) indexed as bare URLs with no way to discover the noindex.
export default function robots(): MetadataRoute.Robots {
return {
rules: { userAgent: "*", allow: "/" },
};
}
```
This file is itself a first-wave audit fix (commit `90c3567`) — before it, the operator console had neither the meta tag nor the robots file. The internal admin console had carried a noindex meta since its first commit but no robots.txt at all, so the second wave gave it the same explicit file (commit `f8acf17`), mirroring the rationale.
Check number 2: for every property you want out of search results, confirm the noindex is _reachable_. Disallow plus noindex is not belt-and-suspenders — the belt hides the suspenders.
## Your product mints URLs on your domain — noindex them at the product level
This is the checklist item that only shows up when your product is itself a website. Clanker Support serves an iframe-able full-page chat shell at `/embed/:key` for hosts that can't run third-party script tags (the details are in [everything that broke while shipping an embeddable widget](https://clankersupport.com/blog/shipping-an-embeddable-widget)). Every customer project gets one of these URLs — on **our** API domain.
Follow that to its bad ending: a customer embeds the iframe, their page links to our URL, a crawler finds it, and now a search for the customer's brand can surface `api.clankersupport.com/embed/` — a bare chat shell wearing their brand color — next to, or instead of, the customer's own site. Nobody involved wants that, and the customer can't fix it, because the page isn't theirs.
Since the shell is served by a Hono route on workerd, not a Next.js page, the fix is a response header (commit `d276077`, in `apps/api/src/routes/embed.ts`):
```ts
// apps/api/src/routes/embed.ts
// Iframe chrome, not content: per-project embed URLs on the api host must
// never appear in search results next to the customer's own site.
c.header("x-robots-tag", "noindex");
```
Check number 3, and the framing we'd push hardest: for a SaaS, index hygiene is a **product decision**, not an SEO chore. Every URL pattern your product generates — embed shells, share links, preview pages, per-tenant subdomains — is a page you are publishing on your customers' behalf. Decide its index status when you design the feature, because by the time it ranks, it's an incident.
## What we ship for AI crawlers, and what we deliberately don't block
The AI half of the audit had two parts: what we add, and what we refuse to subtract.
What we added: `/llms-full.txt`, the [llmstxt.org](https://llmstxt.org) companion to the `/llms.txt` link map we already served. Where `llms.txt` is a table of contents, `llms-full.txt` is the entire text of every blog post in one plain-markdown file, newest first, so an AI system can ingest the content without crawling each page. The builder (`apps/marketing/src/lib/llms-full-txt.ts`) is 55 lines, pure, and unit-tested; the only subtle line rewrites root-relative links, because a markdown link like `/pricing` means nothing once the text leaves our domain:
```ts
// apps/marketing/src/lib/llms-full-txt.ts
// Root-relative markdown links would be resolved against whatever
// domain serves this text, so make them absolute site URLs.
const absolute = p.content.replace(/\]\((\/[^)\s]*)\)/g, `](${siteUrl}$1)`);
```
Honesty about what this buys: llms.txt is an emerging convention, and consumption by the major engines is unproven. Google has said plainly that AI Overviews need no special AI files — ordinary indexable HTML is the input — while other engines are less explicit, and some tooling does fetch these files today. We treat the pair as cheap insurance: one static route, a pure function, six unit tests, and content we'd publish anyway in a format that costs a crawler one request instead of twenty.
What we refused to subtract: the audit checked every robots surface on all five domains for AI-crawler blocks — GPTBot, ClaudeBot, PerplexityBot, Google-Extended — found none, and ratified that as policy rather than leaving it as an accident of defaults. Blocking AI crawlers is a defensible choice for a publisher whose content _is_ the product. Ours isn't; it's an open-source support agent, and the people who might use it increasingly ask an AI assistant what to use. An engine that has read our engineering posts can cite them; one that's blocked at robots.txt recommends whoever wasn't. The honest cost is that AI answers built on your content can substitute for visits to it. For a vendor blog, we'll take citation over control — being the source the answer names is the point of writing.
Check numbers 4 and 5: know your AI-crawler stance instead of inheriting it from a robots.txt someone wrote in 2019, and if you publish llms.txt files, hold them to the same testing standard as any route.
## The smaller line items
Four more findings, one paragraph each, because they'll each cost someone a quiet month.
**The docs subdomain had no baseline at all.** docs.clankersupport.com served no robots.txt, no sitemap, no canonicals, no OG or Twitter cards, and no structured data — the app was two weeks old and every one of those defaults to "missing" (commit `ac538e0` adds all five, plus `TechArticle` and `BreadcrumbList` JSON-LD on every page). Subdomains grow faster than their metadata; audit every host you answer on.
**A brand suffix ate every title.** Appending "— Clanker Support Journal" pushed every blog post title past the ~60-character SERP cutoff, so Google truncates or rewrites them (commit `b433b72` drops it, and adds a dedicated `seoDescription` field because our comparison tldrs and migration intros ran 260–380 characters against a 160-character meta limit). Measure the title as rendered, suffix included.
**Sitemap `lastModified` was a lie.** We stamped build time on every entry, which claims every page changed on every deploy. The comment in `seo.ts` now enforces the rule: only blog entries carry the field, from real publish/update dates, because a timestamp on everything "teaches crawlers to distrust the field."
**The showcase's robots.txt was a 404 page.** Our live-demo site returned HTML for `/robots.txt`, had no `metadataBase` (so its OG image resolved against nothing), and — a Next.js footnote worth knowing — an `opengraph-image` file alone emits only `og:image`; without explicit `openGraph`/`twitter` blocks there's no `og:title` or card type around it (commits `7ad1f32` and `5ad2bf7`, one from each wave).
## What we can't tell you yet
The cross-domain wave merged three days ago, and even the oldest first-wave fixes have had two weeks in production — not enough for Search Console to say anything. So we have zero results data: no ranking movement, no AI citations to report, no before/after chart. Anyone who ships SEO changes on Thursday and reports wins on Sunday is selling something. What we can vouch for today is the mechanism behind each fix: the shallow merge is documented Next.js behavior we verified in prerendered HTML, the Disallow-hides-noindex trap is how Google has worked for years, and the embed-shell noindex closes a real path to ranking against our own customers. We'll report back when Search Console has something worth quoting, including if the answer is "nothing moved."
If the llms.txt items made your own list, the [free llms.txt generator](https://clankersupport.com/tools/llms-txt-generator) we host will build the link-map file from your page list — no sign-up attached. And if you run the nine checks and find your equivalent of the imageless Twitter card, we'd genuinely like to hear what it was: the whole audit started because we grepped our own HTML for a tag we were certain was there.
# We set a cache trap for our own support agent — in our own llms.txt headers
URL: https://clankersupport.com/blog/knowledge-base-recrawl-cdn-cache
Published: 2026-07-20
Category: Engineering
URL knowledge sources are snapshots, and everyone knows snapshots go stale. What we missed is that the refresh itself can be served by any CDN cache between the crawler and the origin — so "Recrawl → success" can silently re-store the pre-deploy content. The trap on our own site was set by our own cache headers.
URL knowledge sources go stale in two layers: they are point-in-time snapshots, so deploying new docs changes nothing until someone recrawls — and the recrawl itself can be answered by a CDN cache sitting between the crawler and the origin, silently re-storing the pre-deploy content. Defeating that second layer takes a unique cache-busting query parameter, `cache: "no-store"`, and no-cache request headers, in that order.
The part that made us wince: the trap on our own site was set by our own hands. Our `llms.txt` — the plain-text index we publish so AI crawlers can read our content — is served with `cache-control: public, max-age=3600`. It is also a URL knowledge source for our own support agent. Click Recrawl within an hour of a deploy and the agent would have refreshed itself with the copy from before the deploy. PR #153 (commit `d905c42`, merged 2026-07-19) closes the hole; this post is the anatomy.
## The staleness layer we already documented
We wrote up layer one in [why our support agent doesn't use RAG](https://clankersupport.com/blog/ai-support-agent-without-rag): a `url` source is fetched once, at creation, and stored as extracted text — at most 200 KB read, 20,000 characters kept, 10-second timeout (`MAX_BYTES`, `MAX_CHARS`, `TIMEOUT_MS` at the top of `apps/api/src/lib/fetch-url.ts`). Refresh is a manual Recrawl button in the dashboard. Deploying new docs does not update the agent until someone clicks it.
We then demonstrated the failure mode on ourselves. We shipped backend SDKs — the pip, gem, and Composer packages — asked our own widget about them, and it had no idea they existed. The diagnosis was boring: nobody had recrawled the source. Layer one, working exactly as (badly) designed — the button fixes it in one click.
But staring at what that button actually does — one plain `fetch` from a worker to a URL — surfaced the layer underneath. The click is not the whole path.
## Why a successful recrawl can still fetch stale content
A `fetch` from your crawler to an origin is not a private conversation. If the URL sits behind a CDN — and docs sites, marketing sites, and anything on modern hosting almost always do — the response can come from an edge cache that stored it earlier, and the cache will keep serving that copy until its `max-age` runs out. Deploying new content does not evict it on its own. Your crawler gets a 200, a plausible body, and no header that screams "this is from before your deploy."
Our refresh endpoint (`apps/api/src/routes/sources.ts`) then does the natural thing: stores the content and stamps `lastFetchedAt` with the current time. The dashboard's status chip reads Ready. Every signal an operator can see says fresh. The bytes are from an hour ago.
That is the expensive property of this bug: it is indistinguishable from success. A failed crawl shows an error and keeps the old snapshot. A cache-poisoned crawl shows the same Ready chip and keeps the old snapshot too — it just launders the timestamp.
## Our own llms.txt would have poisoned our own agent
The concrete instance lives in `apps/marketing/src/app/llms.txt/route.ts`:
```ts
// apps/marketing/src/app/llms.txt/route.ts
return new Response(body, {
headers: {
"content-type": "text/plain; charset=utf-8",
"cache-control": "public, max-age=3600",
},
});
```
That header is correct for its audience. `llms.txt` exists to be hammered by AI crawlers, and an hour of edge caching is basic politeness. But the same file does double duty as a knowledge source for our own agent, and for that consumer the header means: any recrawl within an hour of a deploy re-stores the pre-deploy index, reports success, and re-dates the snapshot.
Being exact about the blast radius: this is a _would have_, not a _did_. We found the hole while fixing the missing-recrawl problem above; we have no evidence a cache-poisoned recrawl ever fired, and no customer got a stale answer we can trace to it. It's a latent footgun we happened to catch while holding it. We wrote the cache header for crawlers in one app and the crawler it defeats in another, and neither file knew about the other until July.
## How to force a fresh fetch through CDN caches
The fix, `fetchFresh` in `apps/api/src/lib/fetch-url.ts`, layers three defenses, strongest first:
```ts
// apps/api/src/lib/fetch-url.ts
const u = new URL(url);
u.searchParams.set("__recrawl", crypto.randomUUID().slice(0, 8));
busted = u.toString();
// …
const res = await fetchBypassingCache(busted, signal);
if (res.ok) return res;
// …
return fetchBypassingCache(url, signal);
```
| Defense | What it does | Why it isn't sufficient alone |
| -------------------------------------------------------------- | ---------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| Unique `__recrawl` query param per crawl | Changes the cache key, forcing a miss on any cache that keys on the full URL | Some origins reject unknown params (signed URLs); a CDN configured to strip query strings ignores it |
| `cache: "no-store"` on the fetch | Bypasses the requesting runtime's own HTTP cache | Governs only your side — no effect on CDNs in the path; older workerd throws on the option |
| `cache-control: no-cache` + `pragma: no-cache` request headers | Asks shared caches to revalidate with the origin | The big CDNs don't honor request-side cache directives by default |
The ordering is the lesson. The headers are what HTTP offers for exactly this situation, and they go last because in practice they do the least: Cloudflare and CloudFront, in their default configurations, ignore a client's `Cache-Control` request header entirely. The query param is a cruder tool — it doesn't ask the cache anything, it makes the cache's own key lookup miss — and that's precisely why it works everywhere caches key on the full URL, which is the default on every major CDN. Each crawl gets a fresh random value, because a stable buster would just get the busted URL cached instead. One of the five tests the fix added in `fetch-url.test.ts` pins exactly that: two crawls of the same URL must carry different `__recrawl` values.
## Why cache: no-store throws on older workerd
Defense two has a runtime problem we've met before. Our API runs on workerd, where [the platform's constraints are build-your-own-adventure](https://clankersupport.com/blog/cloudflare-workers-every-node-sdk-broke), and `cache: "no-store"` is only accepted on compatibility dates from 2024-11-11 onward. Before that, the option doesn't get ignored — it throws:
```ts
// apps/api/src/lib/fetch-url.ts
try {
return await fetch(url, { ...init, cache: "no-store" });
} catch (e) {
if (
e instanceof Error &&
/the 'cache' field|unsupported cache mode/i.test(e.message)
) {
console.warn("[fetch-url] runtime rejected the cache option", {…});
return fetch(url, init);
}
throw e;
}
```
The saving grace is that workerd rejects the option _before any network I/O_ — "The 'cache' field on 'RequestInitializerDict' is not implemented" — so falling back to a plain fetch can never duplicate an in-flight request. The fallback keeps the busted URL and the no-cache headers, dropping only the option the runtime refused. This matters mostly for self-hosters, who run whatever compatibility date their config pins.
The error match is deliberately narrow, and a test guards its edges with a hostile hostname: a DNS failure for `cdn.cachefly.net` contains the word "cache" but must not be misread as a compat rejection and swallowed into a retry.
## Don't let the cache-buster break a working source
An extra query param is not free. Signed URLs — S3 presigned links, anything with a `sig=` — can reject a request whose params don't match the signature. A cache-busting fix that turns a previously-working source into a 403 is a regression wearing a safety vest.
So `fetchFresh` retries: if the busted URL fails for any non-abort reason, it refetches the original URL untouched. That path can still be served stale by an edge cache — which is exactly as bad as before the fix, and no worse. And if even that fails, the refresh endpoint keeps the old snapshot and stamps `lastError` instead of blanking the content. The five tests pin the whole contract: a unique buster per crawl carrying the no-store option and both no-cache headers, the compat fallback, the original-URL retry for signed URLs, the narrow error match, and no retry after the 10-second timeout aborts.
## Your agent is only as fresh as the worst cache in the path
None of this is specific to our no-RAG design. Any knowledge pipeline that ingests over HTTP — full RAG with embeddings, prompt-stuffed snapshots like ours, a nightly scraper feeding a vector store — has its freshness bounded by every cache between the crawler and the truth. The embedding step can't vectorize content the CDN didn't hand over.
The transferable checklist, in the order that earns its keep: bust the cache key with a unique query param, because that defeats caches that ignore your headers; request `no-store` for your own runtime's cache, with a fallback if your runtime predates the option; send `cache-control: no-cache` anyway, for the caches that listen; and never let the buster regress a URL that worked without it. Then go audit your own properties for the recursive case — the content you publish _for_ AI consumers is the content most likely to sit behind a generous `max-age`, and it may well be feeding your own agent.
And treat every "last fetched" timestamp in your pipeline with appropriate suspicion. It dates the fetch. It says nothing about the bytes.
# Changelog: July 2026, so far — the platform month
URL: https://clankersupport.com/blog/llmchat-changelog-july-2026
Published: 2026-07-14
Category: Changelog
Two weeks in and July is already our densest month: an approved WordPress plugin, a Shopify app, agents that take actions, dark mode everywhere, and a smarter, chattier widget.
We usually write these when the month closes. July's first two weeks shipped more than most full months, so here's everything so far — with more to come before August.
**Clanker Support is an approved WordPress plugin.** The plugin went through WordPress.org review and is now [live in the plugin directory](https://wordpress.org/plugins/clanker-support/): install it, paste your project key under Settings → Clanker Support, and the widget is on every page — no code, and it survives theme changes. The [launch post](https://clankersupport.com/blog/wordpress-ai-support-plugin) tells the whole story, including the zip file that was secretly a tar. A 1.0.1 followed within the week.
**A Shopify app, running end to end.** The Shopify app is built and deployed: a zero-permission theme app embed that puts the agent on your storefront without touching your orders, customers, or products. The App Store listing is in Shopify's hands; until it lands, the one-line script tag works on any store today — the [Shopify guide](https://docs.clankersupport.com/integrations/shopify) has both paths.
**The agent can now do things, not just say things.** Agent integrations shipped: the agent can look up a customer's order on Shopify or book a meeting through Cal.com, right inside the conversation, with a "Working on it…" indicator while an action runs. We hardened this layer before shipping it — SSRF guards, per-conversation action limits, and an audit log of every action the agent takes — and scoped the agent to support-only via a base system prompt, so it stays your support agent even when a visitor tries to make it something else.
**Dark mode, everywhere.** The widget now supports `data-theme="light" | "dark" | "auto" | "host"` — auto follows the visitor's OS, and host mirrors your site's own theme toggle live, so the widget flips the instant your page does. The dashboard got dark mode too, and inline embeds accept a theme parameter so a dark page never frames a white chat.
**The widget leads with chat.** Conversations now start in the chat itself — the contact form is opt-in per project, for teams that want a name and email up front. Alongside it: an expandable large panel, admin-defined starter question chips (with a live chat preview in the dashboard while you edit them), an end-of-conversation rating prompt, and a proper "start a new conversation" flow.
**Say "human" and it listens.** If a visitor explicitly asks for a person — "can I talk to a human", "agent please", "I don't want to talk to a bot" — the escalation button appears immediately, before the usual message threshold. The matcher errs toward showing the option: a false positive costs one extra button; a false negative traps a frustrated customer with a bot.
**Quote-reply in the chat.** Visitors can reply to a specific earlier message, so "what about this one?" stays unambiguous in long conversations — for the visitor, the agent, and your team reading the thread later.
**A notification bell in the dashboard.** New conversations, escalations, and new visitor messages across the whole workspace, in one feed — and clicking a notification opens that exact conversation, even if you're already in the inbox.
**An official React / Next.js package.** [`@clankersupport/widget-rsc`](https://www.npmjs.com/package/@clankersupport/widget-rsc) is on npm: one server component in your layout instead of a script tag. There's a [tutorial](https://clankersupport.com/blog/nextjs-ai-support-widget-server-component) if you're on Next.js or any React 19 app.
**A real docs site.** Product docs now live at [docs.clankersupport.com](https://docs.clankersupport.com) — a getting-started path, a page per dashboard surface with real screenshots (light and dark), and integration guides for WordPress, Shopify, and the React SDK.
**Free tools.** An [AI support savings calculator, CSAT calculator, canned response generator, and llms.txt generator](https://clankersupport.com/tools) — free, no signup, built because we kept needing them ourselves.
**Email that behaves.** Mail sent to your team address now forwards into the inbox reliably (and stops bouncing retries), and escalation replies keep threading straight back into the conversation.
That's two weeks. The Shopify listing decision, Slack notifications, and the public usage API are still in flight — see you at the end of the month.
# Add an AI support agent to Next.js, WordPress, Shopify, or any site
URL: https://clankersupport.com/blog/add-ai-support-agent-any-stack
Published: 2026-07-11
Author: Ismail Ghallou
Category: Guides
One hub for every install path we ship: the universal script tag, the React Server Components SDK, the WordPress plugin, the Shopify app embed, and the iframe. Working code for each, plus a candid guide to picking one.
You can add an AI support agent to any website by pasting one `
```
That's the whole install. `widget.js` is a single self-contained file — React, the chat UI, markdown rendering, streaming, all inlined — served with a five-minute cache. The script mounts the widget into a shadow DOM appended to `document.body`, so your site's CSS can't break the widget and the widget's styles can't leak into your page. It loads `async`, so it never blocks your page render.
Your project key is safe to expose in HTML. It only identifies which project answers the chat — clankersupport.com itself runs the widget with its real key committed in the repo. The dashboard generates this snippet pre-filled for you under **Projects → your project → Widget → Install**.
Configuration lives in exactly five `data-*` attributes:
- **`data-project`** (required) — your project's public key. Without it the script throws instead of silently doing nothing.
- **`data-api`** (optional) — the API origin. Defaults to whatever origin served `widget.js`.
- **`data-brand`** (optional) — accent color; defaults to `#111827`. Most people set this in the dashboard instead.
- **`data-mode`** (optional) — `bubble` (default, the floating launcher) or `inline`.
- **`data-escalation-threshold`** (optional) — how many visitor messages before the widget offers a human. The agent default is 3.
There is no sixth attribute. Position, welcome message, starter questions — those are project settings in the dashboard, fetched at runtime, so you can change them without touching your site's HTML.
The `data-api` default is the detail we're most pleased with. Here's the actual resolution logic from the widget source:
```ts
const apiUrl =
script?.dataset.api ??
(script?.src ? new URL(script.src).origin : window.location.origin);
```
The widget derives its API origin from the URL that served the script. So if you [self-host](https://clankersupport.com/blog/the-case-for-self-hostable-ai-support) — the whole product is MIT-licensed, bring your own model keys — you serve `widget.js` from your own domain and the exact same snippet points at your own API. Zero config divergence between hosted and self-hosted.
If you're on Rails, Django, Laravel, plain HTML, Astro, Vue, Hugo — anything that renders a `