Status, 6 October 2026. This page describes production, the bot that has posted hourly since August. Its successor, v0, runs beside it with its own database and its own test chat, and is the design going forward. Production retires once v0 fetches its own articles.

RingFacts · as it works today

Architecture Overview

What the bot is, what it remembers, and what happens in one hour of it running. This page describes the code on main and nothing else — no history, no plans. Where this page and the code disagree, the code is right.

Every section opens with a grey In code strip naming the files that do the work.

1What it does

A small group of people follow three MMA fighters. RingFacts reads the news about those fighters every hour and posts to a private Telegram group the things worth reading — once each, in the group's own language, without the same story arriving five times from five websites.

Almost all of the work is throwing things away. In a typical hour it fetches a few dozen articles and sends nothing, because it has seen them all before or they are not really about anyone on the list. The interesting question is never "can it find news" — RSS does that. It is "can it tell that these nine articles are one thing, and that this tenth one only mentions him in passing."

2The pieces

In code hunter.js the hourly job· server.js the standalone chat responder· setup.sh every deploy command

There are four moving parts and no framework anywhere.

Everything else is plain JavaScript. The RSS parsing, the article-text extraction, the Google News link decoding, the duplicate rules and the message formatting are all hand-written with no libraries. The project has three dependencies: the Postgres driver, the Anthropic SDK, and a small offline language detector (eld) that decides which headlines to translate.

3What it remembers

Two different things, and keeping them apart is the central idea.

A story

One piece of news. One statement, one event, on one day. Every article reporting it points at the same story.

Nine outlets writing up the same press conference are nine articles and one story. The group should hear about it once.

This is what "have we already sent this" means.

A claim

One fact about a fighter, written as a single English sentence — that a fight is booked, that he is injured, that he said something.

A claim has a life: it starts as a rumour and becomes confirmed when an official source says so. It can be supported by many articles over many days.

The difference matters because they answer different questions. A story answers "is this new to the group?". A claim answers "what do we actually know, and how sure are we?". An article can open a story and mint no claim at all — plenty of real news is not a claim about anybody's career.

4The tables

In code schema.sql the tables· lib/db.js every query

Five tables. Every article ever seen is kept, including the junk — the record of what was rejected is what makes it possible to check the rejecting.

An article can be held for several different reasons, and the row records which: story (the news is already told), url (the same page under a different address), embedding (the fallback rule found it too similar to something sent), wrong_subject (not about this person), untrusted_source (a site with a record of keyword spam), tangential (real, but he is only mentioned in passing), and send_failed (Telegram refused the message; it will be retried).

5One run, step by step

In code hunter.js main → huntSubject → classifyItem

A run handles each fighter in turn. For one fighter:

Find articles

Two independent sources, so an outage in one does not blind the bot. Google News is searched once per name alias: English for everyone, plus Ukrainian for the two Ukrainian fighters and Google's Spanish edition for Topuria, so coverage in his own country is not missed. Six publisher feeds (UFC, MMA Fighting, Bloody Elbow, Sherdog, Sport.ua, Marca) are fetched once per run for everybody and then filtered by name. Only articles published in the last 24 hours are considered, and at most five unseen articles per fighter per run.

Drop what we have seen

Anything whose address is already in the database is dropped immediately, before anything expensive happens.

Read the article

Google News does not give out real links — it gives wrapped redirect links that a browser can follow and a script cannot. The bot decodes the wrapper to find the real address, notices when that address turns out to be something already stored, and then fetches the page and pulls the text out of it. This happens before any decision about duplicates, because the thing making the decision needs to read the article.

Turn it into a vector

The headline and the first 1,500 characters of the body go to Gemini's embedding model, which returns a list of numbers representing the meaning. Two articles about the same event end up with similar numbers even in different languages. This is used to shortlist candidates — it is no longer used to make the decision.

Decide what it is

The three closest stories from the last seven days are pulled up and shown to Claude alongside the article, which answers one question: is this article one of these stories, or something new, or a reply to one of them, or not about this fighter at all? This is §6.

Decide whether to send it

If it is new, more questions follow — is it worth a reader's attention, does it contain a claim, is the claim big enough for its own alert. This is §7.

Send, and write everything down

The rows are written first, article by article: the article, any new story, any new claim, and the links between them. Then the messages are built from them and sent. If Telegram refuses a message, the rows that message carried are marked as never-delivered so a later run picks them up.

6The decider

In code lib/matcher.js the prompt and the answer shape· lib/db.js storyShortlist

This is the one place a language model is trusted with a decision, and it is asked a deliberately narrow question. It sees the article's headline, source, publication date and a 1,200-character excerpt of its text, plus up to three existing stories described in one line each. It must answer with a structured form, not free text:

The answer is never trusted as given. Every field is checked against the list of values it is allowed to take, and anything off-menu is downgraded rather than stored. A story id that was never shown to the model is discarded. A claim that does not name the fighter is dropped. Every downgrade goes toward caution — the worst case is the bot posting an article without inventing anything about it.

What it is told about MMA

The prompt carries the rules that are genuinely domain knowledge, and they are specific because vague ones failed. The stages of one fight are different news: the booking, the weigh-in, the result, the bonus, the callout afterwards. An article saying a booked fight happened does not confirm the booking — it reports a new fact. Around a fight, each angle is its own story, down to one story per bookmaker for betting odds. And an article mostly about another fighter that mentions this one is not the wrong subject — it is news with no claim, which readers still see.

7Choosing what to send

In code hunter.js postOutcome· lib/tier.js the prominence rule· lib/untrusted.js the spam rule

An article that survives the decider still has to clear three more things, all of them plain code.

Is the site itself spam?

Some websites stuff fighter names into unrelated content to get indexed. A domain is judged on its own record: at least five articles seen, at least half of them already judged not about the fighter, and never once a readable article body. All three must be true. A ratio alone would silence real outlets whose name-filtered feeds run 35–45% off-target; a blocked fetcher alone would silence real outlets that refuse cloud servers.

Would a reader learn anything?

The decider's yes/no answer to that question folds an article into the quiet pile even when it technically contains a quote. A quote nobody learns from is not news. There is one exception: a booking, a result, an injury or a negotiation is never folded this way, because an event is news whatever a model thinks of the article describing it.

Is he actually in this story?

Two rules decide whether an article gets its own headline or is folded into a quiet list. First, the decider's judgement: if it read the article and called the fighter background colour, that is enough. Second, a counting rule measured against the real archive — if the headline does not name him, and the body is long enough to judge, and it names him once or not at all, it is folded. The two are an OR: neither overrules the other's demotion, because one is a measured rule and the other is a model's opinion, and neither has earned the right to rescue what the other rejects.

Articles folded this way are queued, not deleted. They are written down with their reason and a separate daily job exists to collect them into one quiet list. That job has no schedule firing it, so in practice these articles are recorded and never shown.

8Rumour to confirmed

In code hunter.js recordPost / recordJoin· domain/mma.js officialSource

A claim is born rumor unless the article announcing it came from an official source, or the decider judged that the article reports the promotion's own announcement; then it is born confirmed. For MMA, official means exactly one thing: ufc.com. Record-keeping sites like Sherdog and Tapology are treated as credible media, never as authority.

Later, when another article arrives that the decider says is the same news, and it comes from the official source, and it asserts the claim rather than denying it — only then does the claim flip to confirmed. The bot replies to its own original message with a ✅ Confirmed line, so the correction lands underneath the rumour instead of scrolling away from it.

Denials are recorded and never acted on. An article denying a claim is linked as evidence with a stance of "denies". Nothing in the code flips a claim on a denial — that would let one outlet's pushback erase a fact.

9What the group sees

In code hunter.js assembleMessages → deliver· lib/telegram.js sending

Three kinds of message, in this order.

A ceremony — its own standalone post, reserved for a confirmed fight announcement. This is the only thing that gets a message to itself.

🚨 Fight announced

Daniil Donchenko will fight Sergio Soriano at UFC Paris on 5 September.

— UFC

The digest — one message per fighter per hour, carrying everything else. Rumours about big things get a marked line at the top; everything else is a bullet with the headline, the source as a link, and how long ago it was published.

🔎 Ilia Topuria

🕵️ Rumor: Topuria is in talks to return in December — MMA Fighting, 2h ago

• Gaethje reveals why Topuria was easy to predict — Marca (translated from es), 4h ago

A confirmation — a threaded reply to the message where the rumour first appeared.

Headlines in any language other than English or Ukrainian are translated into English before sending (in practice mostly Spanish); those two are left alone, because the group reads both. For articles from English or Ukrainian feeds the headline itself is checked, offline, since Google's English edition carries Spanish articles too. A failed translation posts the original rather than nothing.

10When things break

In code hunter.js throughout· lib/googlenews.js the circuit breaker

Nothing that fails takes the run down, and nothing is lost from the archive. But the fallbacks are more careful than what they replace, so "it fails open" would be wrong.

Each publisher feed logs how many items it matched and how many it discarded. A name filter that has quietly rotted looks exactly like a quiet news week; these counts are what separate the two. Once a day the three core tables (items, claims, claim_sources) are copied to Google Cloud Storage, because the free Postgres tier keeps six hours of history and nothing more; stories and feedback are not in the copy.

11Configuration

In code domain/mma.js what kind of thing· watchlist.js who

The pipeline knows nothing about MMA. Two files do, and they are separate on purpose.

The domain

What kind of thing is tracked. Which outlets to read, whose word counts as official, the vocabulary of claim types, and the sentences spliced into the model's prompt.

Swap this file and the same machinery tracks musicians instead. domain/example-music.js exists to prove the seam is real, and is clearly marked as never having been run.

The watchlist

Who is tracked. Three fighters, with their name in each language for searching, the surname stems used to filter publisher feeds, and hints about who they might be confused with — a namesake, a relative in the same sport.

Stems, not full names, because Ukrainian declines surnames and a full-name match would miss most sentences.

12Cost

The database, the job, the scheduler and the storage are all inside free tiers. The only real cost is the AI calls, and the expensive one is the decider: one Claude Haiku call per surviving article. The stable part of that prompt — the rules, which are identical for every article about the same fighter and are most of the tokens — is sent as a cached block so it is not paid for again and again.

13What is not measured

The bot can be shown not to repeat itself: of the 54 stories the live decider has opened, 45 reached the group and each arrived once. But that counts the pipeline's own story objects, and if it reads one occasion as two stories then both are posted while the number stays clean. There are 15 pairs of posted articles sitting in different stories while being nearly identical by vector distance, so there is real room for that.

What cannot be said yet is how often it sends the right thing. That needs two numbers: of what it sent, how much was worth sending; and of what was worth sending, how much it sent. The second has no denominator that can be queried — a story the pipeline never recognised has no row to count. It only exists once a person has read the archive and said what was there. That work is underway, and until it lands this project quotes no effectiveness percentage.