Skip to content
AI Agency

The AI Forge · Advanced · F5

Quality Engineering with AI

AI-augmented testing, and testing/evaluating AI systems that don’t pass or fail cleanly

Duration
2 days · 9:00–17:00 each day (13h contact, incl. breaks + lunch)
Participant level
Intermediate: regular AI users
Format
Instructor-led labs, build-along
Participants
8 included · 12 maximum (flexible group size)
Prerequisite
F4 or test-automation experience (or equivalent)
Ecosystem
Agnostic: any model provider or stack
Price
EUR 7,000 (EUR 250 per extra participant)
Certification
Optional ISTQB® CT-GenAI pathway with an accredited partner, in setup · on request

Amounts in EUR, excl. VAT.

A trainer presenting to participants working on laptops in a bright training roomai-generated

What participants will be able to do

  • Generate and maintain test suites with AI assistance
  • Use AI for exploratory testing and bug prediction
  • Integrate AI-assisted QA into CI/CD pipelines
  • Design a full AI-assisted test strategy for your product

Tools & resources

GitHub CopilotPlaywrightPytest (AI-assisted)TestimCI/CD pipeline templates

What changes after this module

Quality engineers use AI to accelerate testing of ordinary software AND build the new probabilistic testing stack that AI systems demand, evals, regression, bias and adversarial checks wired into CI, with methods that outlast today’s tools.

Who should attend

Engineers and data professionals. Intermediate: regular AI users.

Programme

This agenda is indicative. Content, sequencing, and examples are adapted to your team's context, tools, and objectives.

Before you start

  • A laptop with a working development environment and admin rights
  • Proficiency in at least one general-purpose language (Python and/or JavaScript/TypeScript)
  • A code editor/IDE with an approved AI coding assistant enabled
  • An API key for an approved model provider (issued by the company)
  • A Git repository and command-line comfort
  • QA/test-automation or development background
  • A real application under test (traditional and/or an LLM feature) plus its CI pipeline
Day 19:00–10:30Two problems, one discipline: Using AI to test software vs testing AI itself; why probabilistic outputs break pass/fail automation (the “QA crisis”) Lab: Take one flaky/expensive test area and one AI feature: the two systems you’ll carry across both days: and map the right testing approach to each.
10:30–10:45 · ☕ Coffee break
Day 110:45–12:30AI-augmented test automation: AI-generated tests: including derived from requirements and specs: self-healing selectors, coverage analysis and test-data generation for conventional apps; the AI features arriving inside mainstream test tools, and the confidentiality rules for feeding specs and test data to a model Lab: Analyse a real spec with AI to derive the test set, then generate and stabilise an end-to-end suite for that flow.
12:30–13:30 · 🍽 Lunch break
Day 113:30–15:30Testing AI systems: the layers: The three layers of an LLM app (integration/runtime, orchestration, model) and how each fails differently Lab: Instrument an LLM feature and provoke a failure at each layer to see the distinct failure signatures.
15:30–15:45 · ☕ Coffee break
Day 115:45–17:00Lab: eval harness: Build the core evaluation harness for an AI feature Lab: Build an eval harness scoring faithfulness, relevance and safety on a labelled set.
Day 29:00–10:30LLM-as-judge & scoring: Rubric-based scoring, calibrating a judge model, and handling non-determinism with thresholds not equality Lab: Write an LLM-as-judge scorer with a rubric, validate it against human labels, and measure agreement (Cohen’s kappa) before you trust its numbers.
10:30–10:45 · ☕ Coffee break
Day 210:45–12:30Regression, bias & adversarial: Prompt/model-upgrade regression, bias and fairness checks, and adversarial/prompt-injection testing Lab: Add a regression suite that catches a simulated model-upgrade regression, plus a bias and an injection probe.
12:30–13:30 · 🍽 Lunch break
Day 213:30–15:30Quality in CI & production: Evals as CI gates, production sampling and drift, and traceability for audit (EU AI Act / ISO 42001 / NIST) Lab: Wire the eval suite into CI as a quality gate, add a production sampling check, and capture the versioned, timestamped run record that an EU AI Act audit trail requires.
15:30–15:45 · ☕ Coffee break
Day 215:45–17:00Quality strategy & wrap: A tiered testing strategy and what to own vs sample; peer review Lab: Draft the quality strategy for your system and present the CI-gated eval pipeline.

Deliverables

  • AI-assisted end-to-end test suite (conventional app)
  • Eval harness (faithfulness/relevance/safety)
  • Calibrated LLM-as-judge scorer
  • Regression + bias + adversarial checks
  • CI-gated eval pipeline + production sampling
  • Tiered quality strategy

Interested in running this module for your team? Get in touch and we'll tailor the format, dates, and delivery to your context.

Request this training

Your trainer

Has built test and evaluation harnesses for AI systems; quality-engineering depth

Discovery Session

This module requires a full day to deliver a meaningful learning experience. We only offer it in its complete format.

Part of a bigger path

  • EnterpriseEUR 115,000 · up to 30 training days
See the programmes

Frequent questions

Can modules be taken individually?

Yes. Every module stands alone at a fixed price, and every module counts toward a programme if you continue.

Where does training happen?

At your premises or remote, on your dates, for private cohorts. Open sessions run on a fixed monthly calendar.

Which AI tools do you train on?

Yours. Every module ships in four ecosystem editions and runs its exercises on your real stack.

Who delivers?

Inforca's senior consultants and trainers. Flagged modules and the executive track are delivered at senior-expert level.

Discuss this module in a First Call: fit, dates, and the path around it.

Book a First Call

No commitment · our team responds within one business day