QA Automation and Security Testing in One CI Pipeline
Most teams treat test automation and security testing as two separate jobs, run by different people at different times. Security issues then turn up weeks after the code was written, when fixing them is slow and expensive. You don't need a big security team to change that. A QA automation engineer can put the basic security checks into the same pipeline that already runs the functional tests.
This is the layered setup I use. It's four layers, from the cheapest to the most expensive, and each one catches a different kind of problem.
Layer 1: functional automation stays the backbone
Your end-to-end and API tests already log in, create data and call real endpoints. That makes them the best starting point for security checks too, because the hard part (getting an authenticated session in a known state) is done already. A typical Playwright test:
test("user can open their own order", async ({ page }) => {
await page.goto("/orders");
await page.getByRole("link", { name: "Order #1042" }).click();
await expect(page.getByRole("heading", { level: 1 })).toHaveText("Order #1042");
});
Layer 2: security checks written as ordinary tests
Many common security bugs are just missing rules: an endpoint that doesn't check who is calling it, or a header nobody set. They're simple to assert, so write them as normal tests next to your functional ones.
Access control (IDOR / broken object level authorization)
Log in as user A, then ask for an object that belongs to user B. The API must refuse.
test("user A cannot read user B's order", async ({ request }) => {
const res = await request.get("/api/orders/2001", {
headers: { Authorization: `Bearer ${tokens.userA}` },
});
expect([403, 404]).toContain(res.status());
});
Broken access control is first on the OWASP Top 10 (A01:2021) and broken object level authorization is first on the OWASP API Security Top 10 (API1:2023). Scanners rarely catch it, because a scanner can't know which objects belong to which user. Your tests can, because you created that data. I wrote a separate IDOR testing checklist on how to find these by hand first.
Authentication
- Every protected endpoint returns
401with no token, an expired token and a token with a tampered signature. - Logging out actually invalidates the session or token on the server, not only in the browser.
- Password reset tokens can be used only once and expire.
Security headers
test("sends basic security headers", async ({ request }) => {
const res = await request.get("/");
const h = res.headers();
expect(h["content-security-policy"]).toBeTruthy();
expect(h["x-content-type-options"]).toBe("nosniff");
expect(h["strict-transport-security"]).toContain("max-age=");
});
Layer 3: a DAST baseline scan with OWASP ZAP
The OWASP ZAP baseline scan spiders the target and reports passive findings, such as missing headers, insecure cookies and information leaks. It doesn't run active attacks, so it's safe to point at a shared staging environment. In GitHub Actions it's a single Docker command:
zap-baseline:
runs-on: ubuntu-latest
steps:
- name: OWASP ZAP baseline scan
run: |
mkdir -p zap && chmod 777 zap
docker run --rm -v "$PWD/zap:/zap/wrk:rw" ghcr.io/zaproxy/zaproxy:stable \
zap-baseline.py -t https://staging.example.com -r zap-report.html -I
- uses: actions/upload-artifact@v4
if: always()
with:
name: zap-report
path: zap/zap-report.html
-I keeps warnings from failing the build. Start that way, triage the first report with the developers, and only make the scan blocking once the noise is gone. A blocking job that everyone learns to ignore does more harm than a non-blocking one people actually read.
Layer 4: dependency checks
Known-vulnerable libraries are the cheapest risk to catch. Add npm audit --audit-level=high (or pip-audit, or the OWASP Dependency-Check plugin for Maven and Gradle) as its own job, and turn on Dependabot or Renovate so fixes arrive as pull requests.
What still needs a human
Automation covers the known, repeatable checks. These still need manual, exploratory security testing with Burp Suite:
- Business logic flaws, like applying a coupon twice, skipping a payment step or changing a price on the client.
- Chained issues, where two low-severity bugs together become a serious one.
- New features with new trust boundaries, such as file uploads, webhooks, SSO and role changes.
When you find a bug by hand, add an automated test for it. Over time the pipeline turns into a record of every security mistake the product has already made, and it stops any of them from coming back.
A short checklist to start with
- Create two test users with separate data in your test fixtures.
- Add one cross-user access test per sensitive API resource.
- Add authentication tests for missing, expired and tampered tokens.
- Assert the security headers you expect.
- Run a non-blocking ZAP baseline scan against staging and triage it.
- Add a dependency audit job.
- Turn every manually found security bug into a regression test.
None of this replaces a proper penetration test. It does mean the pen tester's time goes on hard problems, not on missing headers and other easy findings.