MediaShelf/README.md
Jess Hallsworth 1ba03dca70
Generalise protected libraries into keep marks
Watch data cannot tell you what is valuable, only what is unwatched. Some
content is held deliberately - home video, 4K copies that are expensive to
re-acquire, shows kept in case someone wants them later - and that judgement
has to be stated by a human and be un-overridable by the score.

One 'keep' concept, applied at four levels: library (a rule, so future
additions inherit it), show, season, movie. Explicit marks below the library
can point either way, so 'keep all of 4K Movies except this one' is
expressible rather than a dead end.

The important detail is identity. Plex reassigns ratingKeys on library
rebuilds and rematches, so marks are keyed on content guid instead -
(library_id, guid) for movies and shows, (library_id, show_guid,
season_number) for seasons. A keep list that silently detaches from its
content would, in v2, mean deleting something explicitly protected. That is
now the single most important test in the suite. Marks are scoped per
library on purpose: Movies and 4K Movies share guids, so an unscoped mark
would protect both copies at once.

Also: kept items hidden from the grid by default with a toggle; a Kept
items screen that surfaces orphaned marks rather than carrying them
silently; the dashboard always shows never-played / kept / available as
three numbers so a growing keep list cannot quietly hollow out the report;
nightly export of keeps and views to JSON, since they are the only data in
the database not reconstructible from Plex and Tautulli; and v2 deletion
refuses kept items ahead of every other guardrail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
2026-09-07 05:38:05 +00:00

89 lines
4.5 KiB
Markdown

# MediaShelf
A self-hosted web app for figuring out which of the thousands of files in a Plex library
are actually worth keeping.
MediaShelf scans a Plex Media Server over its HTTP API and pulls watch history from
Tautulli, builds a local snapshot of every movie and TV season, and joins together facts
that are never shown side by side — date added, size on disk, file path, owning library,
watch count, last watched, and *how much of it anyone actually finished* — into a
sortable, filterable, chartable grid with a tunable **reclaim score** that ranks deletion
candidates.
## Status
**v1 is report-only.** MediaShelf does not delete, move, or modify anything. It produces
a ranked list, saved rule sets, and CSV export. Deletion is designed for in the roadmap
but deliberately not built, so the scanner and the scoring model can be trusted before
anything destructive is wired up.
Nothing is implemented yet — this repository holds the design plus the read-only tools
used to validate it. The design has been checked against the live servers: 65.7 TB across
24 libraries, 2,930 movies and 2,807 TV seasons, 87,640 logged plays from 60 users, and
27.4 TB never played. See `docs/design.md` §2.1.
## What it does
- Full-library ingest from Plex, no agent on the Plex host, no filesystem mounts
- Movies at item level, TV rolled up to **season** level
- Watch data from **Tautulli**, so plays by *every* account are counted and a play that
was abandoned after five minutes is distinguished from one that was finished — Plex's
own history reports those identically. Falls back to Plex session history if Tautulli
is unavailable, and says so rather than degrading silently
- Sort and filter on every metric; charts for size by library, additions over time,
finished vs. abandoned vs. never-opened by size, and size vs. last-watched
- A weighted reclaim score with live sliders, and grace rules so it never recommends
something you added last week
- Named, re-runnable saved views — *"unwatched, older than 2 years, over 10 GB"*,
*"two people started it and nobody finished it"*
- **Keep marks** at library, series, season or movie level, for the things you are holding
on purpose — home video, 4K copies, shows you might want someday. Kept items are hidden
from the working grid and refused outright by deletion. Marks are keyed on Plex content
GUIDs rather than rating keys, so they survive a library rebuild
- CSV export of any view
## Planned stack
Python + Flask, SQLite (WAL), vanilla JS front-end, single container deployed as a
Portainer stack behind Nginx Proxy Manager.
## Roadmap
- **v2** — two-stage quarantine-then-purge deletion, with authentication, a path
allowlist, and an audit log
- **v3** — Emby and Jellyfin support behind the existing `MediaProvider` abstraction
## Validating the design first
Plex and Tautulli are both LAN-only, so `tools/probe.py` exists to check this design
against real data from inside the network. It is **read-only** — GET requests only,
nothing is modified — and has no dependencies beyond the standard library.
```bash
export PLEX_BASE_URL=http://192.168.1.10:32400
export PLEX_TOKEN=...
export TAUTULLI_BASE_URL=http://192.168.1.100:8181
export TAUTULLI_API_KEY=...
python3 tools/probe.py # library shapes, sizes, coverage horizon
python3 tools/probe.py --dump inventory.json # plus a full item inventory
python3 tools/reclaim_preview.py # the actual reclaim numbers
```
`reclaim_preview.py` joins Plex sizes to Tautulli plays the way the application will, and
reports never-played bytes per library — split into the *confident* pool (added while
Tautulli was watching, never played) and the *uncertain* pool (predates Tautulli, might
have been watched). It also prints the largest never-played TV seasons and the
watcher-count distribution used to calibrate the score.
It reports Tautulli's coverage horizon, the finished/abandoned/never-opened split across
real plays, library shapes and sizes, path roots, multi-version items, and whether
Tautulli is actually watching the Plex server you think it is. Credentials are redacted
from all output including error messages, so the result is safe to paste anywhere.
## Documentation
- [`docs/design.md`](docs/design.md) — the full software design
- [`tools/probe.py`](tools/probe.py) — read-only reconnaissance script
- [`tools/reclaim_preview.py`](tools/reclaim_preview.py) — read-only reclaim numbers
- [`tools/mockserver.py`](tools/mockserver.py) — mock Plex/Tautulli for testing the probe