Hercules's Fifth Labor: The Digital Stables
King Augeas received enormous divine herds from the gods. The cattle never got sick, and they multiplied and multiplied. The stables went uncleaned for years. Hercules did it in a single day — he diverted two rivers through them.
We all have a digital stable. It goes by different names: "ARCHIVE 2012," "Old laptop," "Drive D — sort it later." Files pile up on their own too — the messenger auto-saves, screenshots happen at the press of a button, email pulls in attachments, the torrent client downloads things "just in case." They don't get sick, they don't smell, they barely take up space — so they just sit there.
Here's the evidence from real archives:
Over the available connection (~30 Mbit/s, measured) uploading this much data to the cloud would take about 262 days — versus the 96 it actually took to process on-site. Storage would run $84/month, and processing on a comparable GPU would run $150–190/month. And that's without the next step: the data still has to move from storage onto the GPU machine itself for processing — even over a connection three times faster (100 Mbit/s), a batch of a few terabytes takes about 3 days.
1. Who's the likely user
The person with disks in a closet. The label says "IMPORTANT." Inside — several copies of the same thing, a "photos to sort" folder with thousands of files, archives with unknown contents. Can't throw it out, because what if something valuable's in there. Can't sort it, because when?
The marketer in dozens of chats at once. Work chats with clients, topic channels, professional communities. Somewhere in there — live clients, unresolved requests, other people's problems. But the stream is so dense there's no time to read it, and hiring an analyst is expensive.
The student. Years of notes, scanned handouts, lecture and seminar recordings — often in several languages at once. Filed away in folders named things like "3rd year classes" and "need to review." Finding the right fragment before an exam is a whole investigation.
The instructor, school or university. Teaching materials pile up over decades: course notes, lecture recordings, correspondence with students, draft handouts in different languages. The archive grows faster than there's time to review and update it.
The engineer. Datasheets, schematics, calculations, technical documentation — accumulated over years, tens of gigabytes. Can't go to the cloud: confidential, or just doesn't want to feed a stranger's data center with their own knowledge. Half the files have no text recognition, meaning they simply don't exist for search — and somewhere out there, another engineer already solved the exact problem they're stuck on, in a folder just as private as this one. (More on that.)
The musician or songwriter. Concert recordings, drafts, duplicate takes of the same piece in different performances, cassette tapes, digitizations. Knows there's a unique performance or unreleased recording somewhere in there. Doesn't know which of hundreds of files it's in — and doesn't know that someone in the audience that night has the angle their own recording is missing. (More on that.)
The music or literary club. Recordings of concerts, creative evenings, and poetry readings pile up night after night, but there's no one left to sort out who performed and what was actually said or played — the lineup keeps changing, and the archive stays collective, belonging to no one in particular. Often several members recorded the same evening from different seats, and nobody's ever compared notes. (More on that.)
The archivist of a music, poetry, or cultural scene. A self-appointed keeper of recordings from a genre or community's events — a bard-song circle is one example, far from the only one. Concerts, readings, evenings going back decades, held together by whoever still remembers who performed what and when. The archive is only ever as complete as one person's own recordings — even though other attendees, over the years, were very likely recording the same evenings from their own seats.
The archivist. The keeper of a fonds — personal, family, or someone else's archive placed in their care. Media, documents, and recordings from different eras all mixed together, with no single descriptive system. The task isn't to "delete" — it's to understand what's actually there before deciding anything.
2. What the product does
Four tools under one roof — four rivers diverted through four different stables.
Stream one: the library. Takes a folder of files of any type — books, datasheets, schematics, documents — and has a local neural network write an annotation for each one. Prepares unrecognized PDFs for OCR. The output: a proper index you can search by meaning, instead of guessing from a filename like doc_final_FINAL_v3_USE_THIS.pdf.
Stream two: audio and video. Takes recordings — concerts, voice messages, calls, messenger voice bubbles — and turns them into timestamped text. If it's music, it tries to identify exactly what was performed. If it's a conversation or interview with several people, it splits the lines by speaker and summarizes who said what.
Stream three: correspondence. Takes a full export of a channel or chat, collects all the text, downloads voice messages and videos, transcribes them, assembles it into one long document, and summarizes it. The output: who these people are, what they care about, whether there are potential customers among them — plus a ready-made brief for any neural network: here's the audience, here's their language, here's what to offer them.
Stream four: photos. Takes clusters of similar photos — say, every picture from one event or period — and has a local vision-capable neural network describe what's in them: who's in the shot and how many people, where it was taken, what's happening. Not every single photo out of thousands — a representative sample within each cluster, so the archive gets described in a reasonable amount of time instead of weeks of GPU work.
Everything runs locally. Files never leave.
3. What category it belongs to
A local AI scout for personal collections — a new category with no settled name yet: not a cloud transcriber, not corporate document management, not a file search tool, and not a media server, but something in between — with one key difference: it runs on your own hardware, with your own data, and its output isn't a file list — it's a structured library. For every single file — a topic, a language, a document type, key entities, a short summary. This is done ahead of time, without rush or deadline, before anyone urgently needs it — not at the last minute under pressure for a specific task. The finished library can be fed to any neural network, database, or search system — they take it from there, finding what's needed and building the interface; Digercules doesn't replace that next step, it makes it possible in the first place, by taking on the labor-intensive preparation nobody ever has the resources for.
4. Why it beats the closest alternatives
All of these tools — AI file organizers, desktop search, document-management systems with OCR, ordinary transcription — only work once an archive has already been put into machine-readable shape: text recognized, files described, everything sorted. That preparation is exactly what they don't do and can't do — they assume it as a given.
Digercules does exactly that missing part: a mixed archive (text + audio + video + photos + scans) becomes a single structured library, trained on the specifics of the particular collection — entirely locally. From there, any of the tools above, any neural network, or any search system can work with that library — Digercules doesn't compete with them, it gives them the one thing without which they're useless.
5. What tasks it can handle
| Task | Result |
|---|---|
| Figure out what's on the drive | An archive map with annotations for every group of files |
| Remove duplicates and obvious clutter | Deduplication without losing anything that matters |
| Make a library searchable | Annotations + OCR → search by meaning |
| Transcribe hours of recordings | Timestamped text, summarization, identification |
| Understand who's in a chat | An audience portrait + a prompt for the next step |
| Assess what's unique vs. junk | Rare / typical / obsolete / needs attention |
| Make sense of a photo archive | Descriptions from a representative sample within each photo cluster |
6. How it helps
The person with disks in a closet runs the tool, points it at a folder, goes and has tea. A few hours later: here's the clutter — delete it? Here are the family photos. Here are work documents — careful, might be confidential.
The marketer loads in channel exports. Gets back: these are noise, don't waste time. This one — there are potential customers here, here are their pain points, here's a ready-made prompt for the next conversation.
The engineer sees their entire library at once, for the first time: what exists, what's outdated, what can't be found anywhere else. The annotations are written in their own language — because the tool was trained on their own collection.
Students and instructors get a searchable archive of notes, scans, and lecture recordings, even across several languages at once — what used to require flipping through files by hand is now searchable by meaning.
Musicians and songwriters get a transcript and catalog of concert recordings identifying exactly what was performed in each one — including takes and drafts nobody else would ever have relistened to.
A club gets a catalog of its own evenings: who performed, what they read or played, which recording to search for what — instead of an archive held together by the memory of a couple of longtime regulars.
Archivists get, not a file list, but a structured assessment of the entire fonds: what's valuable, what duplicates what's already known, what needs immediate attention.
7. What it costs and what it takes
Free — but only the preliminary estimate. From a description of your archive or a screenshot of the folder, we'll tell you whether it's realistic to sort and roughly how long it would take. The tool itself is paid — a one-time purchase with no cap on volume or on time: you install it on your own machine and sort the whole archive locally, however much data there is. If your hardware can also handle the heavy processing (transcription, identification), the license doesn't stop you from running that yourself either.
When your own hardware isn't enough. Further processing is either rented on an outside GPU machine of your choosing at your own expense (for a market reference, see the figures earlier on this page), or we take on that multi-day coordination ourselves. In that case the price is set by the actual volume found — there's no fixed number or ceiling today; it's a separate arrangement for your archive.
What we need from you. The archive itself (a drive, a folder, a chat export) and an answer to what you want to learn or get out of it.
If we coordinate it. What goes out isn't the archive — it's a compact working export of it (example: 2 TB of video recordings → about 2.7 GB), sent directly to an agreed-upon machine, never to a public cloud.
If something goes wrong. The support window for retrying a failed run is 48–72 hours from when you report it; a re-run is a separate arrangement.
Not yet spelled out on this site: the intake form, a sample contract, exactly where data gets uploaded, the processing schedule and payment schedule, and the acceptance criteria for the result.
There's a separate, emotionally different situation — when the archive isn't your own but came to you from someone else. Covered separately: Inherited Archive.
The same technology doesn't just work for personal archives — if you represent an organization (a company, studio, bureau, clinic, library), see Digercules for Organizations.