AI Data & Knowledge Engineer

innovapptive

Hyderabad 8 Years Exp Posted 66d ago

Job Description

1. Architect the AI Knowledge and Data Layer

  • Design and implement data ingestion and embedding pipelines to convert structured and unstructured content into vectorized representations.
  • Build a unified data schema connecting maintenance, production, and safety data across SAP, Maximo, OSI PI, and SCADA systems.
  • Integrate vector databases (Pinecone, Weaviate, Qdrant, or Chroma) into the AI Platform (MCP) to enable context-aware retrieval.
  • Optimize query efficiency and relevance through hybrid search (semantic + keyword) and metadata tagging.

2. Operationalize RAG (Retrieval-Augmented Generation)

 

  • Implement document chunking, embedding, and retrieval pipelines for PDFs, work orders, shift logs, and incident reports.
  • Develop automated retraining and re-indexing mechanisms to ensure freshness of data.
  • Collaborate with AI Platform Architect to link retrieval flows into agent orchestration layers.
  • Validate precision, recall, and latency metrics for semantic retrieval using real production workloads.

3. Build AI Data Governance and Observability

 

  • Define data lineage, quality metrics, and access control for AI knowledge repositories.
  • Embed telemetry for data latency, embedding drift, and retrieval accuracy into Datadog/Sentry dashboards.
  • Partner with the Chief AI Architect to enforce compliance, explainability, and prompt context versioning standards.

4. Collaborate Across Product and Engineering

 

  • Work with Product Managers and Solution Architects to identify key use cases for AI-driven search and knowledge retrieval.
  • Partner with QA to build automated test frameworks for semantic accuracy and retrieval reliability.
  • Collaborate with industrial data teams to extract and normalize sensor, historian, and SAP data for RAG integration.

5. Drive Continuous Innovation

 

  • Evaluate emerging frameworks for knowledge graphs, embeddings, and contextual caching (e.g., LlamaIndex, LangChain, FAISS).
  • Tune embeddings and hybrid retrieval strategies for domain-specific industrial vocabulary.
    • Mentor developers on data preparation and retrieval design for AI-integrated product features.

Similar Openings for You