About
Authorship is clear. Distribution is not.
Semantic Copyright exists so that rights holders can trace their own work across the internet. We are a small team out of copyright law, machine learning and publishing.
Why
Existing tools only look for the same sentence
Copyright monitoring was built on one assumption for a long time: whoever infringes will copy the text as it stands. That is not what happens now. An article gets rewritten, a study gets summarised and published under another byline, a photograph is cropped and filtered, a recording is re-voiced in another language. Character by character the result is different — and its source is obvious to any reader.
Tools that search for exact matches stay silent in all of those cases. The rights holder finds out by accident, usually months later, once the trail has gone cold and the evidence has scattered across dead links.
How we approach it
We represent a work as meaning rather than as a string of words. Text, images and audio get a semantic fingerprint, and newly published material is compared against it. When similarity crosses a threshold, the finding arrives with its reasoning attached: which parts correspond, how closely, and where the earliest known appearance was.
One thing stays fixed. The system does not decide anything. It prepares the finding; the rights holder and their counsel make the call. A copyright question is never just a similarity score — context, purpose and fair use need human judgement, and that is not a gap we intend to automate away.
Principles
Three things that shape how we work
01
No claim without evidence
Nothing is surfaced without its reasoning. Where the similarity sits, how strong it is and what it was measured against are all written into the file.
02
On the maker's side
We build for the people who produce the work: writers, photographers, publishers, studios and independent creators.
03
Automation does not judge
The system crawls, ranks and documents. Whether a takedown notice goes out is always a human decision.
Roadmap
Where we are
- Q4 2025The idea takes shape and the first technical trials run — semantic matching tested on long-form text.
- H1 2026First working version of the matching engine. A narrow crawler network goes up.
- Q3 2026Closed beta opens with a small group of rights holders.
- Q4 2026Settling the shape of the evidence file and building the dashboard. — we are here
- Q1 2027Dashboard and reporting open to general use. The takedown workflow goes live.
- After thatWider image and audio coverage, API access, multilingual monitoring.
Team
Deliberately small for now
We are trying to hold three things together: people who know copyright law, people who have built search and similarity systems, and people who have worked inside publishing. The team page, company details and references will live here once we launch.
If you work in one of those areas and the problem interests you, we would like to hear from you. There is no open posting — which has never stopped us from meeting the right person early.