Who Wrote This?

A research project, explained simply

Can a computer tell who wrote something, or what kind of page it is?

People leave habits in everything they write: favourite words, punctuation, how they start sentences. Web pages have habits too: a shop looks different from a blog. This project measures how well computers pick up those habits, from simple letter counting to today's AI models, and checks the answers honestly.

Start reading

Two puzzles

Both are sorting problems: give the computer a text, and it has to put it in the right box.

What kind of page is this?

Called genre identification. Is a web page a blog, an online shop, an FAQ, a home page, a discussion forum? Search engines and AI companies use this to organise and filter the web.

  • Blog
  • Shop
  • FAQ
  • Front page
  • Forum
  • Download page

Who wrote this text?

Called authorship attribution. We have a few texts by each suspect and one text of unknown origin. Which suspect wrote it? It helps in forensics, plagiarism checks and literary history.

The hard version: the known texts are about one topic (say, Harry Potter fan fiction) and the mystery text is about another (Star Wars). The computer must follow the writer, not the topic.

How a computer "sees" style

One of the oldest and still strongest tricks: chop the text into tiny overlapping pieces of three characters and count them. Those counts form a fingerprint. Similar fingerprints suggest the same writer.

Fingerprint similarity โ€“

How much the two fingerprints overlap (cosine similarity of the counts).

Text A: most common pieces

    Text B: most common pieces

      A space is shown as ยท. Try editing either text: the fingerprint updates as you type.

      The toolbox, from old to new

      The project compares four families of methods on exactly the same tests.

      1. 1. Counting letter patterns

        Tested

        The fingerprint above, fed to a classic sorting algorithm (a support vector machine). Fast, cheap, runs on any laptop, and surprisingly hard to beat.

      2. 2. Small learned models

        Tested

        Tools like fastText learn which words and word pairs point to each box. They learn only from the examples we give them, nothing else.

      3. 3. Pre-trained language models

        Next

        Models that first "read" a huge part of the internet, then get a short extra lesson on our task. This was the idea behind the original 2019 study.

      4. 4. Large language models (chatbots)

        Planned

        Simply ask an AI model: "Here are texts by five writers. Who wrote this one?" No training at all. Is that enough, and what does it cost?

      Fair testing: no peeking at the exam

      A model is like a student. It studies from examples (training) and then sits an exam on examples it has never seen (testing). If it glimpsed the exam while studying, its score means nothing. That mistake is called data leakage.

      Taking turns: 10-fold cross-validation

      We split the pages into 10 groups. In each round, one group is the exam and the other nine are study material. After 10 rounds every page has been in the exam exactly once, and we average the scores.

      Why this matters here

      The original 2019 study reported a model scoring higher on its exam than on its practice tests. That's a warning sign that it may have seen the exam text while studying. This rebuild enforces a strict rule: every model learns from study material only, and every step is recorded so anyone can re-run it.

      Results so far

      First round: the classic methods (families 1 and 2). Higher is better. Each score is an average over many exam rounds.

      Score by method and dataset (%)

      Web genres: share of pages sorted correctly (accuracy). Authorship: macro-F1, the official score of the PAN-18 competition, which treats every suspect equally.

        Show as a table

        97%

        Counting letter patterns sorts web pages into 7 genres almost perfectly, above the earlier results the 2019 study compared against.

        58

        Authorship across topics is much harder. Even the best classic methods score around 58 out of 100, and worst in Polish.

        7

        Texts per suspect. With so few, fastText learns almost nothing on authorship. That's where pre-trained models could help.

        Check: our copy of the official PAN-18 baseline gives exactly the published scores, which confirms the data and the scoring are right.

        The story behind it

        1. 2019

          The original study

          An undergraduate thesis at the University of the Aegean tried an early pre-trained model (ULMFiT) on both puzzles. It reported big wins on web genres and losses on authorship.

        2. 2026

          A careful second look

          Reviewing it found possible exam-peeking, inconsistent numbers and missing code. So everything is being rebuilt from scratch, with a strict, repeatable test.

        3. Next

          Modern models

          Pre-trained models and chatbots, tested the same way. The goal is a public benchmark and a paper that answers: has AI caught up with letter counting?

        Mini glossary

        Character n-gram
        A tiny slice of text, n characters long. "the" and "he " are 3-grams.
        Training / testing
        Studying from examples, then being examined on new ones.
        Data leakage
        When test material sneaks into training, inflating the score.
        Cross-validation
        Taking turns so every example is tested once, then averaging.
        Accuracy
        The share of answers that are right.
        Macro-F1
        A score that weighs every class (or suspect) equally, so ignoring a rare one is punished.
        Pre-trained model
        A model that learned general language from huge amounts of text before seeing your task.
        PAN
        A long-running research competition on authorship and writing style.