Design and develop advanced conversational AI solutions leveraging Large Language Models (LLMs) and RAG architectures.
Optimize information retrieval pipelines using embeddings, vector search technologies, and reranking models to improve response relevance and accuracy.
Design and implement autonomous AI agents capable of utilizing external tools and services through Model Context Protocol (MCP) environments.
Build robust data ingestion and semantic indexing pipelines for complex document types, including structured technical documentation, tables, and multi-column PDFs.
Implement intelligent chunking and document processing strategies to maximize retrieval effectiveness.
Containerize and deploy AI services using modern cloud-native technologies, ensuring scalability, reliability, and maintainability.
Monitor model performance, costs, latency, and output quality through LLMOps practices and observability platforms.
Analyze prompts, traces, and user feedback to continuously enhance AI system effectiveness and user experience.
Collaborate closely with technical teams and business stakeholders to deliver innovative AI-driven solutions.
Requirements
Proven experience with Python and modern web development frameworks such as FastAPI or Flask for building RESTful APIs.
Strong hands-on experience with LLM orchestration frameworks, including LangChain , LlamaIndex , or direct LLM API integrations.
Deep understanding of Retrieval-Augmented Generation (RAG) architectures and semantic search systems.
Experience working with Vector Databases such as Qdrant , Milvus , Pinecone , Weaviate , or pgvector .
Solid knowledge of Prompt Engineering techniques and AI agent development through function calling and tool integration.
Experience with containerization using Docker and orchestration with Kubernetes .
Familiarity with enterprise cloud environments, particularly Google Cloud Platform (GCP) and/or Microsoft Azure .
Experience managing source control, CI/CD pipelines, and software delivery processes using GitLab .
Hands-on experience with LLM observability and monitoring platforms, particularly Langfuse , including prompt tracking, latency monitoring, token usage analysis, and user feedback collection.