Overall – 8+ | AI Exp – 3+
- Python (primary for AI pipelines, LangChain, asynchronous processing).
- C# / .NET (common across Pearson VUE microservice integrations and Assist Portal backends).
- Azure OpenAI Service, Azure Cognitive Search / AI Search, Azure Container Apps / Azure Kubernetes Service (AKS), Azure App Services.
- Deep understanding of the 2-phase RAG pipeline: document ingestion, chunking strategies, vector embedding generation, similarity retrieval, and context-bounded prompt construction.
- Experience calculating and tuning vector similarity metrics (e.g., Cosine Similarity thresholds) to filter irrelevant queries before hitting LLMs.
- Hands-on proficiency with LangChain, LangGraph, LlamaIndex, or Semantic Kernel.
- Managing multi-workspace prompt templates, system instructions, and dynamic few-shot prompt injection.
- Experience with Azure OpenAI Service (GPT-4o, GPT-4, text-embedding models) and open-source models (Llama, Falcon, Mistral) or multi-model routing (Gemini / Claude).
- Optimizing token usage, latency, temperature/top_p parameters, and fallback/circuit-breaker logic.
- Building performant, asynchronous REST APIs for embedding generation, batch document indexing, chat streaming (SSE/WebSockets), and workspace administration.
Parsing, cleaning, and tokenizing multi-format inputs (PDFs, Word docs, HTML Self-Help Articles, URLs).