Evaluating AWS Textract for PDF Parsing - Table Extraction | Reducto

Customers

Pricing
Reducto leads independent benchmark on structured extraction with Deep Extract
July 1, 2024

Evaluating AWS Textract for PDF Parsing - Table Extraction

We evaluated AWS Textract's ability to parse documents.

Our comprehensive RD-TableBench evaluation included AWS Textract Tables among other market solutions for table extraction. Using our benchmark of 1000 manually annotated complex table images, we tested Textract's capabilities across challenging scenarios that commonly occur in real-world documents.

All data points and outputs are available in the original benchmark blog here.

Overall Accuracy

AWS Textract achieved an 80.9% average table precision score in our evaluation. While this places it among the top performers, it falls significantly short of Reducto's 90.2% accuracy rate. This 9.3 percentage point gap becomes particularly meaningful when processing large volumes of business-critical documents where accuracy directly impacts downstream operations.

AWS Textract vs Alternatives

Our benchmark reveals several key insights about Textract's market position:

1. Performance Hierarchy:

  • Textract (80.9%) ranks third in overall accuracy
  • Trails significantly behind Reducto (90.2%)
  • Performs slightly worse than Azure (82.7%)
  • Substantially outperforms Google Cloud (64.6%)

2. Market Position: Despite AWS's strong cloud presence, Textract demonstrates several limitations:

  • Requires using separate table + layout + forms models to properly parse full documents
  • Can be expensive at scale. Running both Textract Layout and Tables costs 1.9c/page. Running both models with Textract Forms can cost as much as $65,000 to process 1M pages.
  • Shows inconsistent performance with complex table structures

3. Technical Limitations: Our testing revealed several challenges with Textract's approach:

  • Struggles with merged cells spanning multiple rows or columns
  • Limited ability to handle nested table structures
  • Only supports English and a few European languages
  • Difficulty processing tables with unconventional layouts

AWS Textract vs Vision Language Models

While Textract (80.9%) outperforms GPT-4o (76.0%), both solutions demonstrate significant limitations compared to modern approaches. Key observations include:

1. Consistency Tradeoffs:

  • Textract provides more predictable results than VLMs
  • However, its rigid parsing rules often miss complex structural relationships
  • Performance degrades notably with non-standard table formats

2. Feature Limitations:

  • Basic cell detection and text extraction
  • Limited understanding of table hierarchy
  • Minimal ability to adapt to unique table layouts
  • No built-in handling of special cases like footnotes or annotations

Conclusion

AWS Textract's performance in our RD-TableBench evaluation reveals both its strengths and limitations as a table extraction solution. While its 80.9% accuracy rate demonstrates reasonable capability for basic table processing, it falls well short of modern solutions like Reducto (90.2% accuracy) when handling complex real-world scenarios.

The significant accuracy gap becomes particularly relevant for organizations processing large volumes of documents, where even small improvements in accuracy can prevent substantial manual review and correction efforts. For instance, in a dataset of 10,000 tables, Reducto's superior accuracy would mean approximately 930 fewer tables requiring manual intervention compared to Textract.

Organizations should carefully consider whether Textract's limitations align with their document processing needs. While it may suffice for basic table extraction tasks, companies dealing with complex documents containing hierarchical structures, dense information, or unique layouts would benefit substantially from more sophisticated solutions like Reducto that offer significantly higher accuracy and more robust handling of complex scenarios.

CTA patternReducto logo

Make your first API call in minutes.

Reducto logoLLM Center