
Closed
Posted
Paid on delivery
Project Title: Senior Data Acquisition & Document Intelligence Engineer We are looking for a highly experienced Senior Data Acquisition & Document Intelligence Engineer to build, operate, and continuously improve a large-scale production data acquisition and document processing system. Project Type: Long-term / Ongoing Engagement: Full-time Experience Required: 5+ years Location: On-site / Hybrid About the Project We collect large volumes of publicly available documents and structured records from hundreds of external web sources every day. The sources are highly inconsistent and frequently change their website structure, formats, URLs, APIs, and document layouts. A significant portion of the data also comes from scanned PDFs, images, and other difficult-to-process documents. We need someone who can take end-to-end ownership of the entire data acquisition pipeline, including web scraping, crawling, document downloading, OCR, data extraction, structuring, validation, quality assurance, monitoring, and ongoing maintenance. This is not a one-time scraping project. We are looking for someone with proven experience building and maintaining production-grade extraction pipelines at scale. Key Responsibilities 1. Web Scraping & Data Acquisition Build and maintain large-scale crawlers across hundreds of websites and external data sources. Handle dynamic websites, JavaScript rendering, pagination, sessions, cookies, tokens, forms, and file downloads. Work with technologies such as Python, Scrapy, Playwright, Selenium, Puppeteer, Requests/httpx, etc. Implement rate limiting, retries, backoff, and request management. Build incremental crawling and change-detection mechanisms. Detect website/source changes and quickly fix broken extractors. Ensure data is collected completely, accurately, and on schedule. 2. OCR & Document Processing Process large volumes of native PDFs, scanned PDFs, images, and other document formats. Build and optimize OCR pipelines for poor-quality documents. Handle skewed/rotated pages, faded text, stamps, seals, watermarks, handwriting, and complex layouts. Apply image preprocessing such as deskewing, denoising, binarization, cropping, and image enhancement. Extract structured information from complex tables, including merged cells, multi-page tables, borderless tables, and nested headers. Process multilingual documents, including Indian regional languages. Benchmark and select appropriate OCR engines based on document type and accuracy. 3. Data Extraction & Structuring Convert unstructured web/document content into structured datasets. Design and maintain data schemas. Use rules, regex, layout-aware extraction, NLP/NER, and LLM-assisted extraction where appropriate. Normalize dates, amounts, units, names, addresses, identifiers, and other inconsistent values. Implement deduplication and entity resolution. Maintain taxonomies, classification logic, and reference/master data. 4. Data Quality & Validation Build automated quality checks for completeness, accuracy, freshness, consistency, uniqueness, and validity. Implement validation and anomaly-detection mechanisms. Maintain gold-standard datasets and benchmark extraction/OCR accuracy. Create human-in-the-loop review processes for low-confidence results. Reconcile extracted data against the original source. Maintain data lineage and audit trails so records can be traced back to the original source document and extraction run. 5. Production Operations Build and operate reliable, scheduled production pipelines. Use Airflow, Prefect, Dagster, or similar orchestration tools. Implement monitoring, logging, alerting, retries, and failure recovery. Track uptime, throughput, latency, coverage, cost, and data quality. Use Docker and CI/CD. Troubleshoot production issues and prevent recurring failures. Document the system, extraction logic, and operational procedures. Must-Have Skills 5+ years of hands-on experience in production data extraction, web scraping, document processing, or data acquisition. Strong Python programming and production engineering practices. Strong experience with large-scale web scraping/crawling. Experience with Scrapy, Playwright, Selenium, Puppeteer, Requests/httpx, or equivalent. Deep practical OCR experience with Tesseract, PaddleOCR, EasyOCR, Google Document AI, AWS Textract, Azure Document Intelligence, or similar. Strong PDF/document processing experience with PyMuPDF, pdfplumber, pdfminer, Camelot, Tabula, or equivalent. Strong SQL and relational database knowledge. PostgreSQL or equivalent database experience. Experience with Airflow, Prefect, Dagster, or similar workflow orchestration. Experience with AWS and/or GCP. Experience with Docker, CI/CD, logging, monitoring, and alerting. Strong understanding of data quality and automated testing. Experience maintaining and improving production datasets over time. Understanding of responsible and lawful web data collection, including [login to view URL], terms of service, privacy, and applicable data protection requirements. Good to Have OpenCV and image processing NLP / NER LLM-based document extraction Multilingual OCR Indian regional language OCR Entity resolution / record linkage Data lineage and provenance Great Expectations, Soda, Pandera, or similar dbt and dbt tests BigQuery AWS S3, Lambda, Batch GCP Cloud Storage Distributed/batch processing Experience with government, legal, financial, regulatory, or public-sector documents Ideal Candidate We are not looking for someone who has only built basic scraping scripts. The ideal candidate should have experience operating a production system involving: Hundreds of Sources → Crawling → Document Acquisition → OCR → Data Extraction → Structuring → Validation → QA → Production Dataset You should understand common failure points such as website changes, broken selectors, missing documents, OCR errors, duplicate records, incorrect table extraction, silent pipeline failures, source downtime, and data quality degradation—and know how to build systems that detect and recover from these issues. To Apply Please send your CV along with a brief description of: The largest scraping/document extraction pipeline you have personally managed. Number of sources/websites handled. Approximate daily/monthly volume of documents or records. Types of documents processed. Technologies and frameworks used. Your measured OCR/extraction accuracy, if available. How you measured and improved accuracy. One example where a production pipeline failed silently, how you discovered it, and what you changed to prevent it from happening again. Please provide specific numbers, technologies, and examples wherever possible. We are particularly interested in candidates with hands-on experience operating large-scale, continuously running production data acquisition systems.
Project ID: 40678242
15 proposals
Remote project
Active 15 hours ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
15 freelancers are bidding on average ₹25,046 INR for this job

With over 5+ years of hands-on experience in data extraction, web scraping, and production-grade system building, I bring to the table a proven track record of overcoming the exact challenges your project entails. As an AI engineer proficient in Python - your preferred language for implementation alongside other vital technologies like Scrapy, Playwright, Selenium, Puppeteer, Requests/httpx - I have a profound knowledge of working with diverse data sources and their ever-changing structures. What differentiates my work is the successful amalgamation of AI technologies with existing enterprise systems and peripherals -something invaluable for a project like yours. Be it working with Odoo ERP or custom IoT hardware design, I can ensure an AI-powered data acquisition and processing pipeline that seamlessly integrates into your ecosystem. Furthermore, my familiarity with AWS, GCP, and Azure will enable an efficient and scalable deployment that suits your needs. Above all technical skills, what truly sets me apart are my relentless commitment to quality and my capability to deliver under stringent timeframes- two key qualities you seek in the candidate. I've always been appreciated by clients for being able to resolve any issues faced during production promptly and preventing their recurrence. Let's together build a state-of-the-art data system that doesn't just meet your present requirements - but adapts aptly as you scale!
₹25,000 INR in 7 days
6.4
6.4

Hi tenderlabs, I will deliver a production‑grade pipeline that crawls hundreds of websites, downloads and OCRs scanned PDFs, extracts structured data, validates and monitors the flow end‑to‑end. I’ll have a stable version running within three weeks for your budget. Ready to start now with a free sample run. Any priority source to begin with? Waiting for your response in chat! Best Regards.
₹25,687 INR in 3 days
5.5
5.5

I have hands-on experience building Python scraping and document/data extraction pipelines using Scrapy, Playwright, Selenium, BeautifulSoup, OCR/NLP, PostgreSQL, and API integrations, with strong focus on reliability, deduplication, validation, and production monitoring. I can own the full pipeline from source acquisition through structured extraction and QA, and I’m comfortable maintaining extractors as sources change.
₹12,500 INR in 3 days
5.6
5.6

As an accomplished software engineer proficient in Python and well-versed in data management and processing, I believe I am a great fit for the role of Senior Data-Acquisition and Document Intelligence Engineer. With over five years of experience, I've successfully built and maintained large-scale extraction pipelines and orchestrated production systems for various organizations. Dealing with a multitude of data sources and file formats, including scanned PDFs, I have become well accustomed to the complexities involved in retrieving and structuring data completely and accurately. Working with technologies such as Playwright, Selenium, and Scrapy, handling everything from web scraping to OCR pipelines won't be new or challenging to me. Moreover, my ability to employ regex, natural language processing/Named-Entity Recognition, combined with an understanding of rule-based patterns for extraction make me equipped to convert your unstructured content into high-quality structured datasets efficiently. My approach focuses on not just on delivering but also ensuring continuous enhancement and monitoring. You can count on me to bring best practices such as automated data quality checks, anomaly detection & rejection algorithms along with air-tight proofing strategies that guarantee accurate extraction even from difficult documents.
₹12,500 INR in 5 days
4.2
4.2

Hello! I have 5+ years of experience building production data acquisition and document intelligence pipelines using Python, Scrapy, Playwright, OCR, PostgreSQL, Airflow, Docker, AWS/GCP, and automated data-quality checks. I can build and maintain the full pipeline from hundreds of changing sources through crawling, document download, OCR, extraction, normalization, validation, lineage, monitoring, and recovery. I also work with scanned PDFs, complex tables, multilingual OCR, OpenCV preprocessing, and LLM-assisted extraction. I can provide my CV plus exact source counts, document volumes, OCR/extraction accuracy metrics, how accuracy was measured, and a real silent-failure prevention example directly in chat. Timeline: Ongoing / long-term I am available to start immediately and can commit full-time. I will respect your existing standards and suggestions, take ownership of production reliability, and would be glad to work with you long term. Thanks
₹25,000 INR in 7 days
3.7
3.7

I've spent 20+ years in data engineering and applied ML, including document-intelligence pipelines that pull structured fields out of messy PDFs and scans with OCR plus a validation layer, so a senior data-acquisition and document-intelligence build is squarely my kind of work. The hard part is not OCR itself, it is trust: acquiring documents reliably at volume and extracting fields that are correct enough to act on, with a way to catch the cases the model gets wrong. What I'd do: - A robust acquisition layer (APIs, watched folders or scraping) that ingests documents reliably. - OCR plus layout-aware extraction (Tesseract or a transformer model) mapping to your target schema. - A confidence and validation pass that flags low-certainty fields for review instead of guessing. - Clean storage and an API or export so downstream systems get the structured output. You get a pipeline that turns raw documents into reliable structured data, with the uncertain cases surfaced rather than hidden. One thing to confirm: what document types and languages are in scope, and roughly what daily volume should it handle? I can start right away.
₹20,000 INR in 14 days
0.7
0.7

Greetings I am Erinc I am a Web Developer since 2020 knowledgeable in Python and Data Processing. I can confirm I will provide the Senior Data Acquisition & Document Intelligence Engineer you're looking for strictly under your budget. I am available to start right now and regularly on desk everyday. Looking forward to hear from you. Thanks for your consideration.
₹25,000 INR in 7 days
0.6
0.6

With 17+ years of comprehensive experience, my team and I have garnered a deep understanding of and extensive expertise in dealing with the nuances and intricacies of data acquisition, extraction, validation, and structuring. Our skillset is a perfect match for your project's requirements--from web scraping to OCR to data quality management. We have handled various complex challenges that align closely with what your project demands: dealing with dynamically changing websites, inconsistent formats, unstructured content, and OCR for scans or images. Having worked extensively with Python, Scrapy, Selenium, and other relevant technologies in our projects ensures our familiarity and proficiency in executing similar tasks for you. A significant portion of our portfolio comprises large scale solutions designed for high volume data acquisition, just like yours. We have built numerous production-grade extraction pipelines from scratch that continuously maintain their relevance despite frequent changes in source structures—a testament to our ability to build robust, scalable and future-proof systems tailored for challenging scenarios.
₹25,000 INR in 7 days
0.0
0.0

Hi, One thing to clear up first: you've listed the role as on-site/hybrid and full-time. I work remotely from Tunisia. If on-site is a hard requirement, I'm not your candidate and I won't waste your time. If the work can be done remotely with overlapping hours, I'd like to be considered. On the work itself — you described the real problem accurately, which most postings don't: sources change silently and extractors rot. So that's where I'd focus. • Extractors defined declaratively per source, not hardcoded, so a layout change is a config fix rather than a code rewrite. • Breakage detection before the data reaches you: per-source schema and volume assertions on every run, with alerts when field-fill rates or row counts drift outside their normal band. A crawler that returns 200 OK and empty fields is the failure mode that quietly corrupts a dataset for weeks. • Raw artifacts stored immutably before parsing, so any extraction fix can be replayed over historical documents without re-crawling. • OCR engine chosen per document class after benchmarking against a labelled gold set — not one engine for everything. Scanned Indian-language documents and clean native PDFs need different paths. • Prefect or Dagster for orchestration, Docker throughout, human review queue for low-confidence extractions. My background is Python data pipelines, LLM extraction systems and document processing. How many sources are live today, and what breaks most often now?
₹25,000 INR in 7 days
2.6
2.6

⭐⭐⭐⭐⭐Hello The real nightmare isn't building crawlers — it's waking up to find that 30% of your sources silently broke overnight because someone changed their HTML structure, and you have no idea which records are missing or wrong until a downstream team flags it three days later. For a pipeline touching hundreds of sources with OCR-dependent documents, that silent failure risk never goes away — it just gets more expensive the longer it stays hidden. That's why I build pipelines with automated anomaly detection baked into every stage, not bolted on as an afterthought. You get a system that knows when extraction quality drops, when sources go dark, or when OCR confidence falls below threshold — and flags it before your downstream teams notice. I can share exact numbers on volume and accuracy metrics in a conversation — and I'd rather walk through what I've built than dump a generic resume
₹25,000 INR in 7 days
0.0
0.0

Dear Client, At Resonite Technologies, we are thrilled to present our bid for the Senior Data Acquisition & Document Intelligence Engineer position. Our proven team, with over 5 years of hands-on experience in building and maintaining large-scale production data acquisition systems, is well-equipped to meet your requirements. We specialize in developing robust web scraping and document processing pipelines across diverse and dynamic web sources. Our expertise includes Python, Scrapy, and OCR technologies like Tesseract and Google Document AI. We have successfully managed complex production environments, ensuring high data quality through automated validation and monitoring. Our team understands the challenges of inconsistent data sources, and we excel in rapid adaptation to changes in website structures and document formats. We prioritize data accuracy and integrity, implementing comprehensive QA processes and maintaining detailed documentation for traceability. We are confident that our experience aligns perfectly with your project needs, and we look forward to the opportunity to contribute to your ongoing success. Best regards, Karthik B Resonite Technologies
₹55,000 INR in 7 days
0.0
0.0

Hi! Data acquisition and document extraction is exactly what I do - I build Python scrapers and pipelines that pull structured data from websites, PDFs and documents reliably, at scale, and keep running when the sources change. The brief's a little short, so could you drop me a message here in the chat with a bit more detail - which sources or documents, and what data you need out of them? Once I know that, I'll tell you right away how I'd approach it and how long it'll take. I deliver clean, verified data - no junk rows or duplicates - and I'll show you a working sample before you commit. Happy to start today. - Bruno
₹25,000 INR in 7 days
0.0
0.0

New Delhi, India
Member since Aug 27, 2026
₹12500-37500 INR
$8-15 USD / hour
$750-1500 USD
₹12500-37500 INR
₹12500-37500 INR
$10-30 USD
₹4000-6000 INR
₹37500-75000 INR
₹12500-37500 INR
₹1500-12500 INR
min $50 USD / hour
₹750-1250 INR / hour
₹100-400 INR / hour
₹750-1250 INR / hour
₹75000-150000 INR
₹400-750 INR / hour
$250-750 USD
₹12500-37500 INR
₹1500-12500 INR
₹12500-37500 INR
₹12500-37500 INR