What kind of page is this?
Called genre identification. Is a web page a blog, an online shop, an FAQ, a home page, a discussion forum? Search engines and AI companies use this to organise and filter the web.
A research project, explained simply
People leave habits in everything they write: favourite words, punctuation, how they start sentences. Web pages have habits too: a shop looks different from a blog. This project measures how well computers pick up those habits, from simple letter counting to today's AI models, and checks the answers honestly.
Start readingBoth are sorting problems: give the computer a text, and it has to put it in the right box.
Called genre identification. Is a web page a blog, an online shop, an FAQ, a home page, a discussion forum? Search engines and AI companies use this to organise and filter the web.
Called authorship attribution. We have a few texts by each suspect and one text of unknown origin. Which suspect wrote it? It helps in forensics, plagiarism checks and literary history.
The hard version: the known texts are about one topic (say, Harry Potter fan fiction) and the mystery text is about another (Star Wars). The computer must follow the writer, not the topic.
One of the oldest and still strongest tricks: chop the text into tiny overlapping pieces of three characters and count them. Those counts form a fingerprint. Similar fingerprints suggest the same writer.
How much the two fingerprints overlap (cosine similarity of the counts).
A space is shown as ยท. Try editing either text: the fingerprint updates as you type.
The project compares four families of methods on exactly the same tests.
The fingerprint above, fed to a classic sorting algorithm (a support vector machine). Fast, cheap, runs on any laptop, and surprisingly hard to beat.
Tools like fastText learn which words and word pairs point to each box. They learn only from the examples we give them, nothing else.
Models that first "read" a huge part of the internet, then get a short extra lesson on our task. This was the idea behind the original 2019 study.
Simply ask an AI model: "Here are texts by five writers. Who wrote this one?" No training at all. Is that enough, and what does it cost?
A model is like a student. It studies from examples (training) and then sits an exam on examples it has never seen (testing). If it glimpsed the exam while studying, its score means nothing. That mistake is called data leakage.
We split the pages into 10 groups. In each round, one group is the exam and the other nine are study material. After 10 rounds every page has been in the exam exactly once, and we average the scores.
The original 2019 study reported a model scoring higher on its exam than on its practice tests. That's a warning sign that it may have seen the exam text while studying. This rebuild enforces a strict rule: every model learns from study material only, and every step is recorded so anyone can re-run it.
First round: the classic methods (families 1 and 2). Higher is better. Each score is an average over many exam rounds.
Web genres: share of pages sorted correctly (accuracy). Authorship: macro-F1, the official score of the PAN-18 competition, which treats every suspect equally.
97%
Counting letter patterns sorts web pages into 7 genres almost perfectly, above the earlier results the 2019 study compared against.
58
Authorship across topics is much harder. Even the best classic methods score around 58 out of 100, and worst in Polish.
7
Texts per suspect. With so few, fastText learns almost nothing on authorship. That's where pre-trained models could help.
Check: our copy of the official PAN-18 baseline gives exactly the published scores, which confirms the data and the scoring are right.
An undergraduate thesis at the University of the Aegean tried an early pre-trained model (ULMFiT) on both puzzles. It reported big wins on web genres and losses on authorship.
Reviewing it found possible exam-peeking, inconsistent numbers and missing code. So everything is being rebuilt from scratch, with a strict, repeatable test.
Pre-trained models and chatbots, tested the same way. The goal is a public benchmark and a paper that answers: has AI caught up with letter counting?