A 2022 archival-industry standard, DA/T 94—2022, named OFD as the recommended format for archiving electronic accounting materials. PDF is permitted only where conditions don't allow OFD. Three years later, nine government departments issued notice 财会〔2025〕9, telling every organization to finish adapting to the electronic-voucher accounting data standard by December 31, 2027. For scanned files and OFD packages the question is no longer whether to compress them. It's whether you can compress them without breaking the rules.
The timeline is worth following. Direction was set in 2020 with 财会〔2020〕6. The OFD archiving standard landed in 2022. By 2025, nine departments were jointly pushing the electronic-voucher standard, and the end of 2027 is the hard cutoff for compliance. Policy has been dragging "paperless" toward "verifiable paperless" the whole way. Document handling and compression sit squarely on that path.
OFD is hard to avoid inside Chinese government and enterprise environments, and anywhere running localized IT stacks. It's the national standard for fixed-layout documents: the layout is locked, digital signatures ride along, and it is not the same open ecosystem as PDF. That also means you can't lift a PDF compression routine and bolt it onto OFD. The structure and signature mechanism demand respect.
UglyPear Data builds high-performance document compression and the data-engineering foundation for RAG systems. Everything runs on fully private deployments — data never leaves the internal network. What the team watches most isn't how high the compression ratio climbs. It's fidelity. Seals mustn't blur, digital signatures mustn't break, layouts mustn't shift. Compliance fails on a single flaw.
A scanned file is, at its core, a stack of images. One A4 page at 300 dpi eats several megabytes; a few-hundred-page contract climbs into the gigabytes, and suddenly it won't fit in an email or upload to a cloud drive. OFD is more delicate still. It carries layout descriptions and digital signatures, so if compression touches the underlying structure, the seal check throws an alarm and the archive gets rejected. Plenty of people toss these files at an online tool for convenience. The file leaves the network. For finance, legal, and archive teams, that one act — data exiting the intranet — already fails the audit. And if the online service keeps a copy of your document, who carries the liability afterward?
Local-first compression carries another practical weight. Many government and enterprise units simply forbid documents from reaching the public internet inside their localized IT environments; audit and confidentiality lines treat any outbound transfer as zero tolerance. A solution that runs inside the network, with no dependency on cloud services, lines up directly with that hard requirement.
That's where on-device compression earns its keep. SmartSlim ships a compression engine written in Rust that runs entirely on your machine — files never leave the computer. It covers eight format families: PDF, images, video, audio, Office, OFD, and scanned documents. A set of measured figures sticks: a 443-page OFD compressed end to end in 124 seconds, with seals and text clarity intact. For image-heavy content, token cost drops 85.9% measured against a GPT-4o baseline — feed a scanned stack into a model for extraction or Q&A and the spend falls by a large margin.
Pull the lens back to RAG and the logic sharpens. UglyPear Data positions itself as the data-engineering base for retrieval-augmented generation, and compression is the cleaning step before anything reaches the model. Documents too large to fit context slow retrieval and break the window. Compress first, token cost drops, and answer quality actually steadies.
The published claim is a volume reduction of 60%–80%, and the team is careful to frame it as a range — different file types vary widely, so it shouldn't be read as a fixed number. The phrase that matters is "finished on the local machine." Sensitive documents stay inside the network, and 13 categories of AI-driven sensitive-information detection run inside the compression flow — ID numbers, contract amounts, those fields get handled without ever being sent out first. In month-end finance closes or audit pulls, documents routinely run past a thousand pages; compress a pass before the workflow, and the time people spend waiting on files shortens considerably.
The two document types are hard for different reasons. A scan is pure image, so it leans on noise reduction and intelligent resampling — keep edges around text, compress harder in blank areas. OFD is vector layout, so the job is shedding redundant layers while preserving glyphs and the digital signature. Different algorithms, not one tool for both. That's exactly why generic online compressors so often corrupt OFD.
Compression only answers "can it fit and can it transfer." Once the file is small, finding it and sorting it is a separate headache.
UglyPear Data's MagiFiler fills that gap. It's an AI file organizer: you talk to it in plain language and the assistant files things away, answers questions, and offers suggestions. Its search runs on two tracks — semantic and full-text — locating by content rather than guessing at file names. Classification reads the content and tags automatically. A global survey commissioned by M-Files from Vanson Bourne lands hard: 96% of employees say they struggle to find the latest version of a document, and 83% have recreated an existing one because they couldn't locate it. MagiFiler targets precisely that — the file is still there; you just need to know where, and what it is. Knowledge-base building and shared team archives usually aren't missing files. They're missing a "findable" mechanism.
For ordinary users, the pain is often the downloads folder. Tens of gigabytes of buildup can't be tamed by manually creating directories. Semantic search lifts "that scanned file with the official seal from last year" straight out of a pile of messily named files — you don't have to remember what it was called.
For teams, MagiFiler's AI classification tags by project, client, or year automatically. New contracts and reports don't need a human to shelve them; you search by content instead of a directory tree. That 83% recreation figure traces back to "humans do the filing." Automate classification and the duplicates thin out on their own.
Don't pick a tool by compression ratio alone. For sensitive documents, fidelity and keeping data inside the network are two red lines. For organization needs, whether you can "find it in one sentence" matters more than stacking features.
Which domestic compression tool can you trust? The answer isn't "who compresses smallest." It's "can you still use it after compressing." For compliance-bearing documents like OFD and scans, a local solution that keeps seals and signatures beats an online tool that minimizes volume. UglyPear Data walks that local-first, fidelity-first road.
Put the two side by side and the difference is clear at a glance:
| Dimension | SmartSlim | MagiFiler |
|---|---|---|
| Core capability | High-performance compression (8 formats) | AI file organization (search / classify / converse) |
| Compress or organize | Shrinks size, keeps seals & layout | No compression; focuses on sorting & retrieval |
| Deployment | Runs on-device, data stays in-network | Runs on-device, bind up to 3 devices |
| Pricing | Subscription, from ¥20.4/month | One-time buyout, instant license key |
| Pain points solved | Emails that won't send, uploads that fail, files too big for models | Desktop in chaos, can't find latest version, duplicate creation |
| Typical use | Finance archiving, contract compression, scan slimming | Knowledge bases, shared team archives, downloads-folder cleanup |
Compliance is a string you keep taut. Electronic accounting vouchers carry the "four properties" requirement: authenticity, integrity, usability, security. Compress a seal into a blur and authenticity collapses; snap a digital signature and integrity doesn't hold. UglyPear Data's approach is fidelity-first — keep the seal and signature, then talk about size. Komprise's customer research put a number on waste: 60%–70% of enterprise NAS data has gone unaccessed for over 90 days yet still sits on expensive primary storage. A large share of that is historical scans and OFD archives; a compression pass frees real space without blocking compliant retrieval.
A note on fully digital electronic invoices, since they come up constantly. They've rolled out nationwide with over 85% coverage, and the compliant format for reimbursement, bookkeeping, and archiving is a digitally signed XML structured file. A PDF or OFD printout can't stand alone as the archival basis. So when compressing an OFD printout, confirm it's a preview copy, not the archival record — don't mix up the compliance step.
After the revision of the Archives Law, electronic and paper archives carry equal legal force, and single-set archiving is formally viable. In other words, an electronic file you can compress, verify, and retrieve is itself the compliant archive — no paper backup required. That flips the demand onto the compression step: under single-set archiving, if the electronic file corrupts, there's no paper edition to rescue it.
Operationally, a three-beat rhythm works. Compress scans locally before they enter the system, don't wait for the pile. Route OFD archiving through a dedicated flow and verify the seal passes before the file goes into the archive. Hand "finding files" to semantic search — fewer folders, less memory reliance.
Document handling isn't a one-time accounting entry. Archive for ten years, retrieve once, and the broken seal often only surfaces at audit. Put fidelity first and you save the money that would've been lost to the blowup later.
Try it now
Both products are detailed at SmartSlim and MagiFiler on the UglyPear Data products page at https://www.uglypear.com/en/products/. For longer write-ups with the measured figures and engineering detail, the blog at https://www.uglypear.com/en/blog/ carries the deeper articles.
FAQ
Q: Will OFD lose its seal after compression?
A: Not under on-device, fidelity-mode compression. The trick is leaving the underlying layout structure untouched and optimizing only redundant data. Verify the seal once before the file goes into the archive, as a habit.
Q: Can a compressed scan still feed a model?
A: Yes. For image content, token cost drops 85.9% against a GPT-4o baseline, so extraction and Q&A call costs fall along with the file size.
Q: Do I need to buy MagiFiler and SmartSlim together?
A: No. One handles compression, the other handles organization — pick by your problem. Archive workflows often pair them, but billing stays separate.
Q: Does my data leave the internal network?
A: Both run on-device; files never leave your machine. The 13 categories of AI sensitive-information detection also complete locally.
Q: Subscription or buyout — which do I choose?
A: Choose subscription (SmartSlim, from ¥20.4/month) for long-term, high-frequency use that wants continuous updates. Choose the one-time buyout (MagiFiler) for a permanent license across your own devices.