Overall – 8+ | AI Exp – 3+
- Hands-on proficiency with libraries like Ragas, DeepEval, TruLens, Promptfoo, or OpenAI Evals to measure:
- Context Precision & Context Recall (retrieval quality from vector stores).
- Faithfulness / Groundedness (checking against hallucination in bounded context).
- Answer Relevance and Toxicity / Guardrail Compliance.
- Analyzing evaluation datasets, confusion matrices, similarity score distributions, and benchmark trends.
- Computing classification metrics (Precision, Recall, F1, Cosine Similarity thresholds) for retrieval systems.
- Azure OpenAI SDK / OpenAI API and LangChain / LlamaIndex for building automated LLM-as-a-judge scoring scripts and synthetic test dataset generation.
- PyTest / Unittest: Writing automated unit/regression tests for prompt templates and model version upgrades.
- Writing complex SQL queries to extract, curate, and version golden test datasets from relational databases (e.g., PostgreSQL, Azure SQL).
- Querying vector stores and telemetry databases (e.g., pgvector, Splunk/log tables) to evaluate real-world candidate/user query retrieval accuracy and fallback rates.
- Handling structured benchmark datasets and prompt configurations using JSON / JSONL, YAML, and CSV.