Ran the probe and a new reclaim preview against Loki and Tautulli from a LAN host. The measured library is 65.7 TB across 24 libraries with 60 users and 87,640 logged plays. Five things in the design were wrong or missing. - Tautulli's get_library_media_info returns file_size 0 for every show section regardless of section_type, so it cannot cross-check TV sizes. Plex is the only size authority for TV, which is 49 of the 66 TB. - solitude: divisor of 3 confirmed correct (61.7% of watched items have exactly one viewer), but weight raised 0.04 -> 0.10 since on a 60-user server "only one person watched this" is real signal. - rejection: abandonment is 4.9% of plays, not the 30% hoped for. Weight cut 0.12 -> 0.06. Kept because it is decisive when it fires. - Cross-library duplicates promoted from "later candidate" to v1: with 16 movie sections, Movies and 4K Movies routinely hold the same film. - Protected libraries added. Family Videos is 38 GB of irreplaceable home video that is 86% "never played" and scores as a perfect delete target. Also splits the reclaim pool into confident (6.5 TB, added while Tautulli was watching and never played) and uncertain (20.9 TB, predates coverage). 80% of the library predates Tautulli, so pre_history is the majority state, not an edge case. Adds tools/reclaim_preview.py, which produced these numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra |
||
|---|---|---|
| docs | ||
| tools | ||
| .env.example | ||
| .gitignore | ||
| README.md | ||
MediaShelf
A self-hosted web app for figuring out which of the thousands of files in a Plex library are actually worth keeping.
MediaShelf scans a Plex Media Server over its HTTP API and pulls watch history from Tautulli, builds a local snapshot of every movie and TV season, and joins together facts that are never shown side by side — date added, size on disk, file path, owning library, watch count, last watched, and how much of it anyone actually finished — into a sortable, filterable, chartable grid with a tunable reclaim score that ranks deletion candidates.
Status
v1 is report-only. MediaShelf does not delete, move, or modify anything. It produces a ranked list, saved rule sets, and CSV export. Deletion is designed for in the roadmap but deliberately not built, so the scanner and the scoring model can be trusted before anything destructive is wired up.
Nothing is implemented yet — this repository holds the design plus the read-only tools
used to validate it. The design has been checked against the live servers: 65.7 TB across
24 libraries, 2,930 movies and 2,807 TV seasons, 87,640 logged plays from 60 users, and
27.4 TB never played. See docs/design.md §2.1.
What it does
- Full-library ingest from Plex, no agent on the Plex host, no filesystem mounts
- Movies at item level, TV rolled up to season level
- Watch data from Tautulli, so plays by every account are counted and a play that was abandoned after five minutes is distinguished from one that was finished — Plex's own history reports those identically. Falls back to Plex session history if Tautulli is unavailable, and says so rather than degrading silently
- Sort and filter on every metric; charts for size by library, additions over time, finished vs. abandoned vs. never-opened by size, and size vs. last-watched
- A weighted reclaim score with live sliders, and grace rules so it never recommends something you added last week
- Named, re-runnable saved views — "unwatched, older than 2 years, over 10 GB", "two people started it and nobody finished it"
- CSV export of any view
Planned stack
Python + Flask, SQLite (WAL), vanilla JS front-end, single container deployed as a Portainer stack behind Nginx Proxy Manager.
Roadmap
- v2 — two-stage quarantine-then-purge deletion, with authentication, a path allowlist, and an audit log
- v3 — Emby and Jellyfin support behind the existing
MediaProviderabstraction
Validating the design first
Plex and Tautulli are both LAN-only, so tools/probe.py exists to check this design
against real data from inside the network. It is read-only — GET requests only,
nothing is modified — and has no dependencies beyond the standard library.
export PLEX_BASE_URL=http://192.168.1.10:32400
export PLEX_TOKEN=...
export TAUTULLI_BASE_URL=http://192.168.1.100:8181
export TAUTULLI_API_KEY=...
python3 tools/probe.py # library shapes, sizes, coverage horizon
python3 tools/probe.py --dump inventory.json # plus a full item inventory
python3 tools/reclaim_preview.py # the actual reclaim numbers
reclaim_preview.py joins Plex sizes to Tautulli plays the way the application will, and
reports never-played bytes per library — split into the confident pool (added while
Tautulli was watching, never played) and the uncertain pool (predates Tautulli, might
have been watched). It also prints the largest never-played TV seasons and the
watcher-count distribution used to calibrate the score.
It reports Tautulli's coverage horizon, the finished/abandoned/never-opened split across real plays, library shapes and sizes, path roots, multi-version items, and whether Tautulli is actually watching the Plex server you think it is. Credentials are redacted from all output including error messages, so the result is safe to paste anywhere.
Documentation
docs/design.md— the full software designtools/probe.py— read-only reconnaissance scripttools/reclaim_preview.py— read-only reclaim numberstools/mockserver.py— mock Plex/Tautulli for testing the probe