UGLYPEAR AI completes its business upgrade: High-Performance Document Compression × RAG Data Engineering PlatformLearn about the New Business →

Performance Benchmarks

No buzzwords, only measured numbers · all data comes from project test reports, verification welcome

Four sets of core measured data

From P2/P3 end-to-end tests, DeepDoc pipeline measurements and the pipeline test report

124 seconds
443-page OFD fixed-layout document
End-to-end full-flow processing
Pipeline test report
0.924
Three-layer parsing QC total score
(full marks on the structure layer)
P2/P3 e2e test report
64%
Table structure mis-detection converged
Dual-model collaborative layout analysis
DeepDoc pipeline 73-page measurement
85.9%
Image token cost savings
(GPT-4o basis)
Multimodal inference cost measured
Note:The test subject was a real fixed-layout document (a long OFD document with tables and embedded images), and end-to-end means the full chain of upload → parsing → cleaning → redaction → compression → chunking → metadata → QC. Figures vary with document complexity — you are welcome to verify with a test run on your own documents.

Parsing quality: how to read the three-layer scores

Every processed document receives a health-check sheet

443-page OFD document measured details

Structure layer (whether layout blocks, heading hierarchy and reading order are correct) —full marks; after combining the relation layer (table row-column relationships, context continuity) and the content layer (text, OCR accuracy), the overall score is 0.924. Table structure mis-detection was converged through dual-model collaboration 64%

Structure layer score1.00
100%
Overall total score (weighted across three layers)0.924
92.4%

Three-Layer Evaluation System

Going live is only the beginning — continuous health checks keep it reliable

Corpus level: is the raw material good

Gatekeeping before documents enter the knowledge base

Measures the quality of preprocessing itself, scored automatically each time a batch of documents is ingested.

  • Parsing completeness
  • Table structure preservation rate
  • OCR accuracy
  • Redaction coverage
  • Permission hit rate

Retrieval level: is retrieval accurate

Whether knowledge can be found

Quantifies retrieval effectiveness with standard question sets to locate recall problems.

  • HitRate@K / Recall@K
  • MRR (mean reciprocal rank)
  • NDCG@K (ranking quality)

Generation level: are answers correct

The answers end users actually see

Measures the faithfulness and relevance of AI answers, directly tied to business experience.

  • Faithfulness (fidelity to the source)
  • AnswerRelevancy (off-target answer rate)
  • ContextPrecision (context precision)

How evaluation closes the loop: one thumbs-down, and the system finds the cause itself

After a user gives a thumbs-down to an answer, the system automatically traces back the knowledge points the answer cited and attributes the issue to a specific step — was semantics cut by chunking, content mistakenly deleted by cleaning, redaction overdone, or metadata missing? The located result flows back into the evaluation set automatically and is regression-verified automatically at the next strategy change; releases are blocked until metrics pass the regression gate.

Compression performance: numbers from our original craft

The compression engine is the platform's foundation and a direct cost item

Typical compression results

Metric Measured value What it means for customers
Average compression ratio60% (up to 80%)Storage costs cut by more than half
Processing speedIn-memory architecture, 3-5x fasterIn-memory processing across the whole flow, no temporary-directory disk IO
Image quality gatingSSIM ≥ 0.92 / PSNR ≥ 30dBSmaller does not mean blurrier — optimal parameters found automatically
Font deduplicationSaves 30-50% of font spaceAutomatic merging of redundant fonts in Office documents
Multimodal token savings85.9% (GPT-4o basis)Feed compressed images to large models and inference bills drop substantially

View Compression Engine Technical Details →

What numbers will your documents produce?

Book a test run with your real documents and we deliver the report

Request a Test Run