Job Description: QA Lead / AI Test Engineer
Experience
11+ years (min. 3 years leading QA teams; 2+ years testing AI/ML or LLM-based systems)
Full-time (Contract W2 only)
About the Role
We are looking for a QA Lead / AI Test Engineer to own quality strategy across our software products and AI-powered features. You will lead a team of QA and SDET engineers, drive test automation at scale, and define how we validate non-deterministic systems such as ML models and LLM applications. You will work closely with engineering, data science, product, and DevOps to ship reliable, safe, and high-quality releases.
Key Responsibilities: Quality Leadership
- Define and own the end-to-end test strategy, test plans, and quality metrics (defect leakage, escape rate, coverage, MTTR).
- Lead, mentor, and grow a team of QA engineers and SDETs; run hiring, performance reviews, and upskilling.
- Act as the quality gatekeeper for releases: go/no-go decisions, risk assessment, and release readiness reporting.
- Partner with product and engineering in requirement reviews, shift-left practices, and sprint planning.
Test Automation & Engineering
- Design and maintain scalable automation frameworks for UI, API, mobile, and performance testing.
- Integrate automated suites into CI/CD pipelines with quality gates, parallel execution, and reporting.
- Drive contract testing, service virtualization, and test data management strategies.
- Champion code quality in test code through reviews, design patterns, and reusable libraries.
AI/ML & LLM Testing
- Build evaluation frameworks for ML models and LLM-based features (chatbots, RAG pipelines, copilots, agents).
- Design test approaches for non-deterministic outputs: golden datasets, semantic similarity, LLM-as-judge, and human-in-the-loop evaluation.
- Test for hallucination, bias, toxicity, prompt injection, jailbreaks, data leakage, and robustness.
- Validate model performance (accuracy, precision/recall, F1, latency, cost per request), drift, and regression across model or prompt versions.
- Verify data quality and pipelines feeding training, fine-tuning, and retrieval systems.
- Use AI tools to improve QA productivity: test case generation, self-healing tests, log analysis, and defect triage.
Performance, Security & Reliability
- Lead performance, load, and resilience testing (including AI inference latency and throughput).
- Collaborate with security teams on OWASP Top 10, OWASP LLM Top 10, and vulnerability testing.
- Support production monitoring, observability, and post-release validation.
Governance & Process
- Establish QA standards, documentation, and best practices across teams.
- Ensure compliance with relevant standards (e.g., ISO 27001, SOC 2, GDPR, responsible AI guidelines).
- Report quality KPIs and risks to senior leadership.
Required Qualifications
11+ years in software QA/testing, with at least 3 years in a lead or manager role.
Strong hands-on experience with automation tools such as Selenium, Playwright, Cypress, Appium, REST Assured, or Postman/Newman.
Proficiency in at least one language: Python, Java, or JavaScript/TypeScript.
Solid experience testing REST/GraphQL APIs, microservices, and event-driven systems.
Experience with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, Azure DevOps) and containers (Docker, Kubernetes).
Working knowledge of performance testing tools (JMeter, k6, Gatling, or Locust).
Hands-on experience testing AI/ML or LLM-based applications, including evaluation metrics and dataset-driven testing.
Strong understanding of Agile/Scrum, SDLC/STLC, and test management tools (Jira, TestRail, Xray, Zephyr).
Excellent communication, stakeholder management, and people leadership skills.
Preferred Qualifications
Experience with LLM evaluation tools and frameworks such as DeepEval, Ragas, promptfoo, LangSmith, TruLens, or Giskard.
Familiarity with LangChain, LlamaIndex, vector databases, and RAG architectures.
Knowledge of ML concepts: supervised and unsupervised learning, model drift, feature pipelines, and MLOps (MLflow, Kubeflow, SageMaker, Azure ML, or Vertex AI).
Experience with cloud platforms (AWS, Azure, or Google Cloud Platform).
Exposure to red teaming and adversarial testing of AI systems.
Experience with SQL and NoSQL databases and data validation tools (Great Expectations, dbt tests).
ISTQB Advanced/Expert certification or an AI/ML certification.
Domain experience in [BFSI / Healthcare / E-commerce / SaaS, etc.].
Core Competencies
Strategic thinking with strong hands-on technical depth
Risk-based testing mindset
Data-driven decision making
Curiosity about emerging AI testing practices
Ability to influence without authority