Turn millions of documents into Spark DataFrames
StabRise builds tools that read PDFs, scans and DICOM files at cluster scale. Extract text, tables and entities, and redact what has to stay private, all on your own infrastructure.
Runs on Apache Spark, Databricks, AWS, Azure and Google Cloud. Built for HIPAA and GDPR workloads.
Projects
Open-source libraries for Spark clusters and for the browser, and a free tool anyone can use online.
Open source libraries

Spark PDF
A powerful, open-source data source for processing and handling PDF files in Apache Spark. Designed to efficiently manage large PDF files with minimal memory usage and scalable performance.
- Open-source
- Supports large files
- Optimized for performance

ScaleDP
An Open-Source Library for Processing Documents using AI/ML in Apache Spark.
- Open-source
- Highly scalable

ScaleDP-TS
ScaleDP for TypeScript. Render PDFs, run OCR, detect signatures and faces, and extract entities right in the browser with WebAssembly or WebGPU, so documents never leave the user's machine.
- Open-source
- Runs in the browser
- WebAssembly and WebGPU
- Same stages as ScaleDP
Free online tool
What teams build with it
Invoice processing
Automate and streamline invoice data extraction to improve accuracy and speed in financial processing.
Clinical trials and medical records
Extract critical data from clinical trials and medical records to enhance research and healthcare workflows.
Anonymization for data science
Ensure data privacy by anonymizing sensitive data for use in data science projects and machine learning models.
Data sharing
Safely share data while maintaining privacy through de-identification and anonymization techniques.
RAG for PDF documents
Build Retrieval-Augmented Generation (RAG) systems for processing large volumes of PDF documents effectively.
Synthetic PII generation
Generate synthetic Personally Identifiable Information (PII) to replace removed or anonymized data.
Why StabRise
- Large files
- PDFs up to 10,000 pages and DICOM files up to 3 GB are processed with minimal memory use.
- Scalability
- Run on a Spark cluster or as a REST API service, on your own isolated servers or on AWS, Azure and Databricks.
- Compliance
- Built for workloads that fall under HIPAA, GDPR and other privacy regulations.
- Security
- Data is encrypted and handled under strict protection protocols at every step of processing.
- Quality control
- Results for each document or page are checked with generative AI and human review.
- Expertise
- Compliance specialists working with data scientists and ML engineers who have spent years on document processing.
Have a document pipeline that needs to scale?
Tell us about your files and volumes, and we will suggest a setup that fits.

