When No One Knows What's In There

Who's the likely user

There's a particular category of people who've inherited someone else's things. A collector died, or a researcher, an archivist, simply a person with a history — and their disks, hard drives, flash drives, external storage ended up with whoever took on the responsibility of preserving them. Sometimes it's a professional archivist. Sometimes a relative. Sometimes a trusted friend.

What they all have in common: they're holding someone else's legacy and don't know what's inside.

The professional archivist — a keeper of oral-history collections, sheet music archives, other collections. Drives from people who are no longer alive arrive regularly. They know they're obligated to go through it. They don't know when or how — because manually, this could take years.

The family of a researcher or scholar — a father spent his whole life collecting material on his subject, died, left behind terabytes. The children want to preserve it, but don't understand what's in there or who might need it.

A cultural institution — a library, museum, or foundation that received a donated collection. Formally accepted. In practice — a box of drives sits on a shelf waiting for a turn that never comes.

What the product does

The starting point is zero. No description, no table of contents, no person left who remembers what was put where. Just files.

First step — quick reconnaissance. In half a minute, the program scans the entire mass and produces a first map: what file types, what languages, what time periods, whether there's any internal structure at all or it's just one big pile.

Second step — careful deduplication. Inherited archives almost always contain copies: the same files in different folders, backups of backups. The program finds exact content-based duplicates and removes the excess carefully: it keeps the copy from the more reliable source, moves the rest to a separate folder — never deletes. Before any action — a preview: look first, act second.

Third step — layer-by-layer analysis. Documents and books are recognized and annotated by a local neural network. PDFs and DJVUs with no text layer are prepared for OCR. Audio and video are transcribed into text with timecodes on every phrase — the program handles old, noisy, or vinyl recordings too, where an ordinary speech detector goes silent: loudness is normalized to the BS.1770-3 standard before recognition.

Fourth step — identification. For audio recordings: who's performing and what — by matching the transcribed text against a database of known works. For video: the program doesn't go through it frame by frame — it finds points of real shot change and offers face recognition on representative frames at those points; a static recording where the camera barely moves yields a handful of frames for this, not thousands. The program doesn't decide on its own who's in a photo or frame — it shows a similarity score and a reference image to compare against, and a person always confirms; a face no one recognizes stays flagged and waits for someone who actually knew these people, rather than being passed off as identified. A name confirmed once is recognized by the system across the entire archive without asking again, and no decision is ever overwritten — only appended to, so the history of the review can always be traced. For every recording, you see what was performed and who's in frame, at what recognition quality, and what remains unidentified.

Fifth step — assessment and report. What's unique here and found nowhere else. What's typical. What's personal and not meant for other eyes. The output: PASSPORT.md — a structured archive passport, suitable for handoff to an institution or to heirs.

Everything runs locally. Files never leave.

What category it belongs to

This is a tool for initial archival reconnaissance — what used to be done only by a person who spent months on manual review. Existing archive-management systems assume you already know what you have. Here, the situation is different: you don't know anything. The program closes exactly that gap — between a box of drives and a legible map of a legacy.

Why it beats the alternatives

Path one: by hand. A person sits down and goes through it file by file. At a volume of several terabytes, this can take years. Most archives stay untouched — not out of indifference, but because there's no resource for it.

Path two: hire a specialist. Expensive, slow, and the specialist ends up doing the same manual work anyway.

The program offers a third path: automatic reconnaissance in a few days of machine time, after which a person sees the full map and makes decisions knowingly, not blindly.

The fundamental difference from cloud services: someone else's legacy is an especially delicate thing. It may contain personal correspondence, unpublished manuscripts, restricted-access documents. Sending it to a corporate cloud is unacceptable. Everything stays in place.

What tasks it can handle

Task Result
A first map of the archive In half a minute — types, languages, periods, overall structure
Find and remove duplicates Careful deduplication: excess goes to a separate folder, not the trash
Recognize documents OCR for books and scans with no text layer
Transcribe recordings Text with timecodes on every phrase
Handle old recordings Loudness normalization, support for noisy and vinyl recordings
Assess recording quality Classification: usable / partial / unrecognizable
Identify who's performing Voice identification of performers
Identify what's being performed Matching against a database of known works
Recognize who's in frame on video Face recognition on representative frames, no frame-by-frame review
Flag what's personal and private What can't be shared or published without permission
Prepare an archive passport A document for handoff to an institution or heirs

How it helps

The archivist gets, in a few days, what would otherwise have taken years: a full picture of someone else's legacy. They see what's unique and needs careful preservation, what can be digitized and opened up, what's personal and should stay closed. The drives stop being dead weight — there's already a finished document in hand, ready to pass on.

A scholar's family understands, for the first time, exactly what their father spent his life collecting — and can hand it to those who need it: a specialized archive, a university, colleagues. With understanding, not guesswork.

A cultural institution can finally answer the question that's been hanging since the donation arrived: what's here, what's unique, what to do with it.


The generation of people who spent their lives collecting, recording, preserving — is passing. With them goes the knowledge of exactly what they collected. Drives age. Formats become obsolete. The window in which this can still be saved and understood isn't infinite.

If the outcome of this reconnaissance is a decision to hand the archive to a library, museum, or foundation, the receiving institution faces the same task from the other side: Donated Collections on digercules.org.