The first real scan logged 15 "episode has no season; skipped" warnings. They
are Firefly S1 in TV Show Archive: Plex returns those episodes with
grandparentRatingKey and parentIndex set and parentGuid present, but
parentRatingKey null. Requiring parentRatingKey meant the entire season was
silently absent from the report - exactly the kind of quiet omission a reclaim
tool must not have.
The season key is now synthesized from show + season number when Plex omits it,
which is stable across scans. Keep marks are unaffected either way since they
key on GUIDs, not rating keys.
Also drops the multi_part flag from ordinary seasons. A season has one part per
episode, so part_count > 1 is normal there and the badge appeared on every TV
row; it now means what it says - a movie held more than once, or a season with
more files than episodes.
Both cases are in the fake server now, so the suite covers them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
history_coverage is only written at the end of an ingest, but
has_completion_data was reading from it. During the first scan the coverage
table is empty while watch_event already holds tens of thousands of Tautulli
rows, so the dashboard announced a fallback that had not happened and claimed
the rejection component was disabled when it was not.
The flag now comes from the events themselves — does any row carry a
percent_complete — which is true the moment Tautulli rows land and false for
Plex-only history. history_source falls back to the scan record when coverage
is absent, so it reads "tautulli" mid-scan instead of null.
Also: the banner named Plex as the source without checking, and now reports
whichever source is actually active, and stays quiet while a scan is running
since the counts are still moving.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
Running on Nox as Portainer stack 113, host port 8086 (8085 and most of 808x
were already taken). Built in place by a git stack — no registry on this
network and no SSH account on Nox — so §11.2 now documents that as the route
taken, with registry and image-upload kept as alternatives. Not yet behind NPM.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
No registry exists yet and there is no SSH onto Nox, so Portainer clones this
repo there and compose builds the image on the box. Adds build: . and moves the
default host port to 8086, which is free on Nox (808x is crowded).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
The application the design describes: Flask + SQLite, Plex for library data,
Tautulli for watch history, report-only.
Structure follows the design's seams. providers/ splits MediaProvider from
HistoryProvider, because on this network library data and watch data live on
different machines and Jellyfin later will have no Tautulli equivalent.
scoring.py implements the reclaim score twice - as a SQL expression for the
live grid (weights change on every slider drag, so storing it would mean
rewriting thousands of rows per drag) and in Python for CSV export and tests,
with a property test over 500 generated rows asserting the two agree.
rules.py compiles saved views to parameterized SQL through a field/operator
whitelist; nothing user-supplied is ever interpolated.
Three properties are enforced by test rather than asserted in prose:
- Ingest is idempotent. Three consecutive full scans leave every count and
every byte total unchanged. A scanner that double-counts produces a report
that looks plausible and is wrong.
- Keep marks survive Plex reassigning every rating key in the library. They
are keyed on content GUID, scoped per library so the Movies and 4K Movies
copies of the same film mark independently.
- Every config variable the app reads is declared in docker-compose.yml, so
a variable set in Portainer can never silently do nothing.
Also found and fixed while verifying against a fake Plex+Tautulli pair:
executescript() commits the pending transaction, so migrations needed their
BEGIN/COMMIT inside the script; replaceChildren() renders null as the literal
text "null"; a hash-only URL change does not reload the document, so deep
links needed a hashchange listener; and SQLite ROUND rounds half away from
zero where Python rounds half to even.
73 tests, no live server required.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
The keep_all capability stays at library level, but no library is marked by
default. KEEP_ALL_LIBRARIES defaults to empty rather than naming Family
Videos, and the design says plainly that every keep in the system got there
because a person put it there.
A keep the application invented would be indistinguishable in the database
from one the user made, which undermines the audit trail the Kept view
exists to provide.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
Watch data cannot tell you what is valuable, only what is unwatched. Some
content is held deliberately - home video, 4K copies that are expensive to
re-acquire, shows kept in case someone wants them later - and that judgement
has to be stated by a human and be un-overridable by the score.
One 'keep' concept, applied at four levels: library (a rule, so future
additions inherit it), show, season, movie. Explicit marks below the library
can point either way, so 'keep all of 4K Movies except this one' is
expressible rather than a dead end.
The important detail is identity. Plex reassigns ratingKeys on library
rebuilds and rematches, so marks are keyed on content guid instead -
(library_id, guid) for movies and shows, (library_id, show_guid,
season_number) for seasons. A keep list that silently detaches from its
content would, in v2, mean deleting something explicitly protected. That is
now the single most important test in the suite. Marks are scoped per
library on purpose: Movies and 4K Movies share guids, so an unscoped mark
would protect both copies at once.
Also: kept items hidden from the grid by default with a toggle; a Kept
items screen that surfaces orphaned marks rather than carrying them
silently; the dashboard always shows never-played / kept / available as
three numbers so a growing keep list cannot quietly hollow out the report;
nightly export of keeps and views to JSON, since they are the only data in
the database not reconstructible from Plex and Tautulli; and v2 deletion
refuses kept items ahead of every other guardrail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
Ran the probe and a new reclaim preview against Loki and Tautulli from a LAN
host. The measured library is 65.7 TB across 24 libraries with 60 users and
87,640 logged plays. Five things in the design were wrong or missing.
- Tautulli's get_library_media_info returns file_size 0 for every show
section regardless of section_type, so it cannot cross-check TV sizes.
Plex is the only size authority for TV, which is 49 of the 66 TB.
- solitude: divisor of 3 confirmed correct (61.7% of watched items have
exactly one viewer), but weight raised 0.04 -> 0.10 since on a 60-user
server "only one person watched this" is real signal.
- rejection: abandonment is 4.9% of plays, not the 30% hoped for. Weight
cut 0.12 -> 0.06. Kept because it is decisive when it fires.
- Cross-library duplicates promoted from "later candidate" to v1: with 16
movie sections, Movies and 4K Movies routinely hold the same film.
- Protected libraries added. Family Videos is 38 GB of irreplaceable home
video that is 86% "never played" and scores as a perfect delete target.
Also splits the reclaim pool into confident (6.5 TB, added while Tautulli
was watching and never played) and uncertain (20.9 TB, predates coverage).
80% of the library predates Tautulli, so pre_history is the majority state,
not an edge case.
Adds tools/reclaim_preview.py, which produced these numbers.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
Tautulli stays local-network-only, so nothing off the LAN can check this
design against real data. tools/probe.py closes that gap from the inside:
GET requests only, standard library only, credentials redacted from all
output including error messages.
It answers open questions 1-4 in one run — Tautulli's coverage horizon,
the finished/partial/abandoned split across real plays, whether successive
plays are being grouped, library shapes and sizes, multi-version items,
and path roots by size. It also runs the pms_identifier cross-check from
section 4.11.
tools/mockserver.py mocks both APIs so the probe is testable without a live
server. Verified against it: the happy path, Tautulli absent, Tautulli
unreachable, Tautulli erroring, Plex unreachable, and an identifier
mismatch. Credential redaction confirmed in every error path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
Tautulli is running at 192.168.1.100:8181, so watch data now comes from
get_history rather than Plex's session history. This matters beyond
convenience: Plex records that a play started, Tautulli records how far it
got. A film three people abandoned after five minutes and a film nobody ever
opened are opposite signals for a deletion decision, and Plex reports them
identically.
- Adds sections 4.8-4.11: Tautulli ingest, completion semantics, fallback,
coverage horizon, and the pms_identifier cross-check
- Splits MediaProvider into MediaProvider + HistoryProvider so library data
and watch data can come from different servers
- watch_event gains percent_complete, disposition, source, and session
merging; media_item gains partial/abandoned counts and pre_history
- New 'rejection' score component, with weights renormalizing when a
component is unavailable rather than silently scoring everything lower
- Deployment: image built off-box and pushed to Nox's Portainer, versioned
tags rather than :latest
- Closes three open questions, opens two smaller ones
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra
Software design for a Plex library analytics and reclaim-reporting tool.
v1 is report-only: no deletion, no filesystem access, Plex API as the sole
data source. Movies at item level, TV rolled up to season level.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVbG48GAXfCZatcmX123Ra