Senior Python Data Engineer (OCR & Document Processing)- remote at Inetum


Company Logo

Inetum is Hiring

Job Info:
  • Company Inetum
  • Position Senior Python Data Engineer (OCR & Document Processing)- remote
  • Location Bucharest, Romania
  • Source SmartRecruiters
  • Published September 16, 2026
  • Category Development
  • Type Full-Time
  • Experience Senior


Job Description

Mission

Design, build, and optimize scalable data ingestion and document processing solutions that transform large volumes of unstructured insurance data into structured, AI-ready information. Enable downstream AI and retrieval systems by leveraging OCR, document intelligence, vector databases, and cloud-native data pipelines.

Responsibilities:

  • Design and implement scalable data ingestion pipelines for processing high volumes of unstructured documents, including PDFs, scans, emails, and Office files.
  • Integrate, configure, and optimize OCR and document extraction technologies to maximize text extraction accuracy and document understanding.
  • Build automated workflows for document parsing, text cleaning, normalization, semantic chunking, and metadata enrichment.
  • Develop connectors and integrations for document sources such as SharePoint, email systems, and enterprise repositories.
  • Design and maintain vector database schemas and retrieval mechanisms to support Retrieval-Augmented Generation (RAG) solutions and AI applications.
  • Ensure document processing pipelines meet enterprise security, compliance, performance, and availability requirements.
  • Implement monitoring, validation, and quality-control mechanisms to identify and manage low-confidence OCR and extraction results.
  • Optimize data processing workflows for scalability, reliability, and low-latency operations.
  • Collaborate with AI Engineers, Backend Engineers, and Platform teams to deliver end-to-end AI-powered document processing solutions.
  • Develop and maintain cloud-native data ingestion solutions on public cloud platforms.

Profile

Professional Experience

  • 5-10 years of experience in Data Engineering, Data Processing, Document Intelligence, or related fields.
  • Proven experience building scalable data ingestion and processing pipelines.
  • Experience working with large volumes of unstructured and semi-structured data.
  • Experience designing cloud-based data solutions.

Technical Skills

  • Strong programming skills in Python.
  • Strong SQL knowledge.
  • Hands-on experience with AWS services, including:
    • S3
    • Step Functions
    • CloudWatch
  • Experience processing unstructured documents such as:
    • PDF
    • Word
    • Excel
    • PowerPoint
    • Email content
  • Experience building connectors and integrations with enterprise content repositories (e.g., SharePoint).
  • Experience with OCR and document extraction tools (AWS Textract or equivalent).
  • Experience designing and implementing data ingestion and transformation pipelines.
  • Familiarity with vector databases and Retrieval-Augmented Generation (RAG) concepts.
  • Experience with software development best practices:
    • Git
    • CI/CD
    • Automated testing

Nice to Have

  • Experience with Vector Databases.
  • Experience with RAG architectures and AI/LLM-based applications.
  • Experience with Azure cloud services.
  • Experience with Databricks.
  • Experience in Insurance, Banking, or other regulated industries.

Benefits

  • Full access to foreign language learning platform
  • Personalized access to tech learning platforms
  • Tailored workshops and trainings to sustain your growth
  • Medical insurance
  • Meal tickets
  • Monthly budget to allocate on flexible benefit platform
  • Access to 7 Card services
  • Wellbeing activities and gatherings