The AI Forge · Advanced · F5
Quality Engineering with AI
AI-augmented testing, and testing/evaluating AI systems that don’t pass or fail cleanly
- Duration
- 2 days · 9:00–17:00 each day (13h contact, incl. breaks + lunch)
- Participant level
- Intermediate: regular AI users
- Format
- Instructor-led labs, build-along
- Participants
- 8 included · 12 maximum (flexible group size)
- Prerequisite
- F4 or test-automation experience (or equivalent)
- Ecosystem
- Agnostic: any model provider or stack
- Price
- EUR 7,000 (EUR 250 per extra participant)
- Certification
- Optional ISTQB® CT-GenAI pathway with an accredited partner, in setup · on request
Amounts in EUR, excl. VAT.
ai-generatedWhat participants will be able to do
- Generate and maintain test suites with AI assistance
- Use AI for exploratory testing and bug prediction
- Integrate AI-assisted QA into CI/CD pipelines
- Design a full AI-assisted test strategy for your product
Tools & resources
What changes after this module
Quality engineers use AI to accelerate testing of ordinary software AND build the new probabilistic testing stack that AI systems demand, evals, regression, bias and adversarial checks wired into CI, with methods that outlast today’s tools.
Who should attend
Engineers and data professionals. Intermediate: regular AI users.
Programme
This agenda is indicative. Content, sequencing, and examples are adapted to your team's context, tools, and objectives.
Before you start
- A laptop with a working development environment and admin rights
- Proficiency in at least one general-purpose language (Python and/or JavaScript/TypeScript)
- A code editor/IDE with an approved AI coding assistant enabled
- An API key for an approved model provider (issued by the company)
- A Git repository and command-line comfort
- QA/test-automation or development background
- A real application under test (traditional and/or an LLM feature) plus its CI pipeline
| Day 1 | 9:00–10:30 | Two problems, one discipline: Using AI to test software vs testing AI itself; why probabilistic outputs break pass/fail automation (the “QA crisis”) Lab: Take one flaky/expensive test area and one AI feature: the two systems you’ll carry across both days: and map the right testing approach to each. |
| 10:30–10:45 · ☕ Coffee break | ||
| Day 1 | 10:45–12:30 | AI-augmented test automation: AI-generated tests: including derived from requirements and specs: self-healing selectors, coverage analysis and test-data generation for conventional apps; the AI features arriving inside mainstream test tools, and the confidentiality rules for feeding specs and test data to a model Lab: Analyse a real spec with AI to derive the test set, then generate and stabilise an end-to-end suite for that flow. |
| 12:30–13:30 · 🍽 Lunch break | ||
| Day 1 | 13:30–15:30 | Testing AI systems: the layers: The three layers of an LLM app (integration/runtime, orchestration, model) and how each fails differently Lab: Instrument an LLM feature and provoke a failure at each layer to see the distinct failure signatures. |
| 15:30–15:45 · ☕ Coffee break | ||
| Day 1 | 15:45–17:00 | Lab: eval harness: Build the core evaluation harness for an AI feature Lab: Build an eval harness scoring faithfulness, relevance and safety on a labelled set. |
| Day 2 | 9:00–10:30 | LLM-as-judge & scoring: Rubric-based scoring, calibrating a judge model, and handling non-determinism with thresholds not equality Lab: Write an LLM-as-judge scorer with a rubric, validate it against human labels, and measure agreement (Cohen’s kappa) before you trust its numbers. |
| 10:30–10:45 · ☕ Coffee break | ||
| Day 2 | 10:45–12:30 | Regression, bias & adversarial: Prompt/model-upgrade regression, bias and fairness checks, and adversarial/prompt-injection testing Lab: Add a regression suite that catches a simulated model-upgrade regression, plus a bias and an injection probe. |
| 12:30–13:30 · 🍽 Lunch break | ||
| Day 2 | 13:30–15:30 | Quality in CI & production: Evals as CI gates, production sampling and drift, and traceability for audit (EU AI Act / ISO 42001 / NIST) Lab: Wire the eval suite into CI as a quality gate, add a production sampling check, and capture the versioned, timestamped run record that an EU AI Act audit trail requires. |
| 15:30–15:45 · ☕ Coffee break | ||
| Day 2 | 15:45–17:00 | Quality strategy & wrap: A tiered testing strategy and what to own vs sample; peer review Lab: Draft the quality strategy for your system and present the CI-gated eval pipeline. |
Deliverables
- AI-assisted end-to-end test suite (conventional app)
- Eval harness (faithfulness/relevance/safety)
- Calibrated LLM-as-judge scorer
- Regression + bias + adversarial checks
- CI-gated eval pipeline + production sampling
- Tiered quality strategy
Interested in running this module for your team? Get in touch and we'll tailor the format, dates, and delivery to your context.
Request this trainingYour trainer
Has built test and evaluation harnesses for AI systems; quality-engineering depth
Discovery Session
This module requires a full day to deliver a meaningful learning experience. We only offer it in its complete format.
Frequent questions
Can modules be taken individually?
Yes. Every module stands alone at a fixed price, and every module counts toward a programme if you continue.
Where does training happen?
At your premises or remote, on your dates, for private cohorts. Open sessions run on a fixed monthly calendar.
Which AI tools do you train on?
Yours. Every module ships in four ecosystem editions and runs its exercises on your real stack.
Who delivers?
Inforca's senior consultants and trainers. Flagged modules and the executive track are delivered at senior-expert level.
Discuss this module in a First Call: fit, dates, and the path around it.
No commitment · our team responds within one business day