Remote Jobs RockRemote Jobs Rock

Senior Python Data Engineer (OCR & Document Processing)- remote

๐Ÿ•’ 5 days ago
Data EngineeringData ProcessingDocument IntelligencePython

๐Ÿ“œ Description

  • Design and implement scalable data ingestion pipelines for processing high volumes of unstructured documents, including PDFs, scans, emails, and Office files.
  • Integrate, configure, and optimize OCR and document extraction technologies to maximize text extraction accuracy and document understanding.
  • Build automated workflows for document parsing, text cleaning, normalization, semantic chunking, and metadata enrichment.
  • Develop connectors and integrations for document sources such as SharePoint, email systems, and enterprise repositories.
  • Design and maintain vector database schemas and retrieval mechanisms to support Retrieval-Augmented Generation (RAG) solutions and AI applications.
  • Ensure document processing pipelines meet enterprise security, compliance, performance, and availability requirements.

๐Ÿ› ๏ธ Requirements

  • 5-10 years of experience in Data Engineering, Data Processing, Document Intelligence, or related fields.
  • Proven experience building scalable data ingestion and processing pipelines.
  • Experience working with large volumes of unstructured and semi-structured data.
  • Experience designing cloud-based data solutions.
  • Strong programming skills in Python.
  • Strong SQL knowledge.
  • Hands-on experience with AWS services, including:S3Step FunctionsCloudWatch
  • Experience processing unstructured documents such as:PDFWordExcelPowerPointEmail content
  • Experience building connectors and integrations with enterprise content repositories (e.g., SharePoint).
  • Experience with OCR and document extraction tools (AWS Textract or equivalent).

โœจ Benefits

  • Full access to foreign language learning platform
  • Personalized access to tech learning platforms
  • Tailored workshops and trainings to sustain your growth
  • Medical insurance
  • Meal tickets
  • Monthly budget to allocate on flexible benefit platform
  • Access to 7 Card services
  • Wellbeing activities and gatherings
Full job description

Company Description

Inetum is a European leader in digital services. For businesses, public sector organizations, and society as a whole, the groupโ€™s 26,000 consultants and specialists work every day to create tangible digital impact: solutions that contribute to performance, innovation, and the common good.

With a presence in 19 countries, working closely with local communities, and alongside its major software developer partners, Inetum supports organizations in their digital transformation challenges with proximity, flexibility, and responsibility. Driven by its purpose, Inetum champions a vision of technology that is useful and well-managed, capable of unlocking the full potential of organizations and society: โ€œLetโ€™s make tech right.โ€

In 2025, the group generated revenue of 2.2 billion euros.

More information at: www.inetum.com

Job Description

Mission

Design, build, and optimize scalable data ingestion and document processing solutions that transform large volumes of unstructured insurance data into structured, AI-ready information. Enable downstream AI and retrieval systems by leveraging OCR, document intelligence, vector databases, and cloud-native data pipelines.

Responsibilities:

  • Design and implement scalable data ingestion pipelines for processing high volumes of unstructured documents, including PDFs, scans, emails, and Office files.
  • Integrate, configure, and optimize OCR and document extraction technologies to maximize text extraction accuracy and document understanding.
  • Build automated workflows for document parsing, text cleaning, normalization, semantic chunking, and metadata enrichment.
  • Develop connectors and integrations for document sources such as SharePoint, email systems, and enterprise repositories.
  • Design and maintain vector database schemas and retrieval mechanisms to support Retrieval-Augmented Generation (RAG) solutions and AI applications.
  • Ensure document processing pipelines meet enterprise security, compliance, performance, and availability requirements.
  • Implement monitoring, validation, and quality-control mechanisms to identify and manage low-confidence OCR and extraction results.
  • Optimize data processing workflows for scalability, reliability, and low-latency operations.
  • Collaborate with AI Engineers, Backend Engineers, and Platform teams to deliver end-to-end AI-powered document processing solutions.
  • Develop and maintain cloud-native data ingestion solutions on public cloud platforms.

Qualifications

Profile

Professional Experience

  • 5-10 years of experience in Data Engineering, Data Processing, Document Intelligence, or related fields.
  • Proven experience building scalable data ingestion and processing pipelines.
  • Experience working with large volumes of unstructured and semi-structured data.
  • Experience designing cloud-based data solutions.

Technical Skills

  • Strong programming skills in Python.
  • Strong SQL knowledge.
  • Hands-on experience with AWS services, including:
    • S3
    • Step Functions
    • CloudWatch
  • Experience processing unstructured documents such as:
    • PDF
    • Word
    • Excel
    • PowerPoint
    • Email content
  • Experience building connectors and integrations with enterprise content repositories (e.g., SharePoint).
  • Experience with OCR and document extraction tools (AWS Textract or equivalent).
  • Experience designing and implementing data ingestion and transformation pipelines.
  • Familiarity with vector databases and Retrieval-Augmented Generation (RAG) concepts.
  • Experience with software development best practices:
    • Git
    • CI/CD
    • Automated testing

Nice to Have

  • Experience with Vector Databases.
  • Experience with RAG architectures and AI/LLM-based applications.
  • Experience with Azure cloud services.
  • Experience with Databricks.
  • Experience in Insurance, Banking, or other regulated industries.

Additional Information

Benefits

  • Full access to foreign language learning platform
  • Personalized access to tech learning platforms
  • Tailored workshops and trainings to sustain your growth
  • Medical insurance
  • Meal tickets
  • Monthly budget to allocate on flexible benefit platform
  • Access to 7 Card services
  • Wellbeing activities and gatherings
Salesloft

Senior Data Engineer

๐Ÿ•’ 3 days ago
Salesloft๐Ÿ‘ฅ 501 - 1000 employees๐Ÿข Computer Software

As a Senior Data Engineer, you will design and build scalable data pipelines, ensuring data quality and governance while supporting security analytics and internal reporting.

Data EngineeringData PipelinesSQLPython
OpenTable

Senior Data Engineer (AI/ML)

๐Ÿ•’ 3 days ago
OpenTable๐Ÿ‘ฅ 1001 - 5000 employees๐Ÿข Hospitality

This role combines modern data engineering with Generative AI, focusing on designing scalable data platforms and pipelines while building production-grade solutions using LLMs and AI agents.

Data EngineeringGenerative AILLMsRAG Architectures
Nagarro1

Senior Data Engineer (Rust)

๐Ÿ•’ 4 days ago
Nagarro1๐Ÿ‘ฅ 10,000+ employees๐Ÿข Information Technology And Services

Lead the architecture and scaling of high-throughput data services and ingestion pipelines while building ultra-low-latency data infrastructure using Rust.

RustData EngineeringAPI DesignCloud Computing

Trusted by Remote Workers