Key Responsibilities
- Build retrieval and verifications over data systems, with shown queries and citations for every answer
- Stand up self-hosted open-weight models serving and embeddings inside each bank’s environment or shared environments; evolve RAG to a dedicated standard
- Design the MCP tool layer that exposes a small, audited set of read-only tools (metrics, documents, customer 360), eventually growing into read/write tools with heavy amounts of regulated, highly sensitive data
- Build and maintain the evaluation harness — golden-question regression, groundedness and retrieval metrics, explicit “I don’t know” behavior — and make it a release gate
- Implement LLM guardrails: PII redaction in prompts and context, prompt-injection defenses, and cost and row limits aligned to regulatory security expectations
- Partner with data teams so the model selects governed metrics from the semantic layer rather than improvising SQL
- Document model architecture, evaluation methodology, and guardrail controls to support customer security reviews and audit readiness
- Track latency, cost, and quality trade-offs across model versions and deployment configurations
Core Competencies
- Accuracy and evaluation orientation — a demonstrated focus on verifiability and groundedness, not just compelling demos
- Production LLM/RAG engineering: retrieval pipelines, tool orchestration, prompt engineering, and guardrail implementation
- Security and compliance mindset: PII handling, prompt-injection defense, and least-privilege tool access aligned to NIST CSF 2.0 principles
- Cross-functional collaboration with data and platform engineering to deliver a governed, auditable AI system
Key Performance Indicators (KPIs)
- Golden-question accuracy — maintained or improved release over release against the verified question set
- Groundedness rate: percentage of assistant answers fully supported by retrieved context
- PII redaction coverage and zero prompt-injection incidents in production
- Model serving latency and cost per query within defined targets
- Evaluation harness adoption as a release gate — zero releases without passing the regression suite
Qualifications
To perform this job successfully, an individual must be able to perform each essential duty satisfactorily. The requirements listed below are representative of the knowledge, skill, and/or ability required.
- 6–10+ years building software, with 2–3+ years shipping production LLM, RAG, or NLP systems used by real people — not prototypes
- A demonstrated focus on accuracy and evaluation, not just demos
- Strong Python and solid software-engineering fundamentals
- Comfort operating self-hosted open-weight models and reasoning about latency, cost, and quality trade-offs
Core Technologies
- Languages: Python
- Serving & inference: vLLM, Ollama; GPU / CUDA familiarity, NVIDIA Enterprise (NVAIE)
- RAG & retrieval: LlamaIndex or Haystack; Qdrant, pgvector; embeddings
- Orchestration: MCP, tool / function calling
- Structured querying: text-to-SQL; semantic layers (Cube / dbt MetricFlow)
- Evaluation & guardrails: groundedness and eval frameworks, PII redaction, prompt-injection defense
Nice to Have
- Experience in regulated or high-stakes domains where a wrong answer is costly
- Fine-tuning, adapters, and retrieval-quality optimization
- Familiarity with banking and finance terminologyƒ
Education and/or Experience
- Bachelor’s degree in computer science, mathematics, or a related technical field, or equivalent hands-on experience
- Experience in the financial services industry or a regulated, high-accuracy AI application environment strongly preferred