Security engineering · 14 min read
A security scanner that files its own pull requests
I built an Azure DevOps task for a large industrial company that does more than fail a pipeline. It detects the stack, runs several scanners, asks an LLM to triage the noise, and can open a remediation PR instead of leaving developers with a screenshot and a headache.
The core idea was simple: a security pipeline should not stop at classification. Most teams already have scanners. What they lack is the layer between raw findings and an actionable change set. So I built a task that wraps multiple tools, normalises their output, enriches it with an LLM pass, and then uses automation very selectively: dependency bumps where the fix is mechanical, targeted configuration edits where a validator can prove the result is still syntactically sound, and a pull request for a human to review before anything reaches production.
The interesting part was not making one scanner work. It was deciding where automation should stop. A system that can edit infrastructure or dependency manifests in your CI environment is only useful if it is boringly constrained. Most of the architecture is really about those constraints.
The shape of the system
The packaged deliverable is an Azure DevOps task extension, but the real engine is a Python package inside it. The task entry point is a small TypeScript wrapper that runs on the build agent, installs the bundled Python dependencies, installs scanner CLIs if needed, and then shells into the Python CLI.
- task: AiaSecurityScan@0
inputs:
azureOpenAIEndpoint: $(AZURE_OPENAI_ENDPOINT)
azureOpenAIDeployment: gpt-4o
failOnSeverity: high
autoFixPR: true
imageTag: ""
skipTriage: falseThat split turned out to be useful. The Azure DevOps integration stays small and versionable as a VSIX, while the Python layer owns the parts that actually change quickly: stack detection, scanner adapters, suppression logic, triage, reporting, and remediation.
| layer | responsibility | why it lives there |
|---|---|---|
| TypeScript task | Collect task inputs, install tools, invoke CLI | Fits Azure DevOps task model cleanly |
| Python core | Detection, scanning, triage, reports, PR logic | Faster to evolve than task packaging |
| Azure OpenAI | Structured triage and targeted fixes | Adds judgment where scanners are noisy |
| Azure DevOps REST API | Create remediation PRs | Keeps review in the existing workflow |
What it actually scans
The repo detector looks for familiar marker files - pyproject.toml, requirements.txt, package.json, Dockerfile, *.bicep, *.tf, and a few others - and builds a simple stack list from that. The scanner orchestration is deliberately less ambitious than the detector.
| scanner | when it runs | what it contributes |
|---|---|---|
| pip-audit | Python requirements-based projects | Dependency CVEs and fixable version bumps |
| TruffleHog | Always | Secret detection across the repository |
| Trivy | Docker repos or explicit image tag | Container and filesystem vulnerabilities |
| Checkov | Bicep or Terraform present | IaC and Dockerfile policy findings |
One detail I like because it is honest: the detector already recognises Node, .NET and C++, but the current scanner wiring is still strongest on Python, containers and IaC. That is better than pretending the coverage is broader than it is. The architecture is ready to expand; the shipped scanner set is narrower.
Important nuance
Normalising scanner output into one model
The core package maps every tool into the same internal Finding shape: scanner, ID, title, description, severity, affected target, fix version, references, and a few LLM-enriched fields. That sounds unglamorous, but it is the reason the rest of the system works. Without a shared model, you cannot sort findings by severity across tools, generate one report, or feed the whole batch into a single triage pass.
Finding( scanner="trivy", id="CVE-2026-xxxx", severity="high", affected="package@1.2.3", fix_version="1.2.4", auto_fixable=true )
The task also supports a repository-level suppression file. If a team has a known false positive or a risk they have consciously accepted, they can declare that in .aia-security-scan.toml instead of playing whack-a-mole in every pipeline run. Suppressed findings still show up in reports; they just stop driving the fail gate.
Where the LLM adds value
I did not use the model as a replacement for scanners. I used it as the layer that scanners are bad at: context-sensitive triage. The task sends the consolidated findings to Azure OpenAI with a strict JSON schema and asks for six things per finding: corrected severity, exploitability, recommended action, suggested fix, whether it is auto-fixable, and whether it looks like a suppressible false positive.
{
"id": "CVE-2026-xxxx",
"severity": "high",
"exploitability": "likely",
"recommended_action": "Upgrade to the patched version",
"suggested_fix": "Pin package >= 1.2.4",
"auto_fixable": true,
"suppress_reason": null
}Two implementation choices matter here. First, the call runs at temperature=0 with structured output rather than free-form prose. I wanted consistent fields, not eloquence. Second, authentication can use an API key or fall back to managed identity via DefaultAzureCredential, which makes the task viable in locked-down enterprise environments without baking secrets into the extension.
This is also where most of the false-positive reduction happens. A secret detector can only say “this string resembles a secret”. A triage pass can say “this is an obvious test fixture” or “this one deserves escalation immediately”. The model is not treated as ground truth, but it is useful as a second-stage filter on top of deterministic scanners.
The remediation paths are intentionally uneven
I did not build one universal auto-fixer. There are two very different remediation modes because the risk profile is different.
1. Dependency bumps
For Python dependency findings, the task takes the boring route on purpose. If pip-audit already knows the safe version, the system runs pip-audit --fix. For some Trivy-reported Python findings, it can patch pyproject.toml or requirements*.txt directly to raise the minimum version. That is a mechanical change, so the automation budget can be higher.
2. IaC and Dockerfile fixes
This is where the genuinely agentic part starts. The newer core package includes a fixer loop built on Microsoft Agent Framework. For each eligible finding it exposes three tools to the model: read_file, write_file, and run_validation. The model has to read the target file, write back a complete replacement, and then prove the result passes a validator.
read_file(path) write_file(path, full_updated_content, explanation) run_validation(path) -> ok | error
Validation is language-aware. Bicep goes through az bicep build, Dockerfiles use hadolint when available, Python uses py_compile, and JSON/YAML are parsed structurally. If validation fails, the change is rolled back. If the model never produces a valid edit, the finding is skipped and left for a person.
How the loop knows when to stop
This is the part I cared about most. The fix loop is not allowed to wander. It is bounded in four different ways.
- It only attempts fixes for a narrow scanner set - currently Checkov findings in the agentic path.
- The tool surface is tiny: read one file, overwrite one file, validate one file.
write_fileandrun_validationhave capped invocation counts, so a bad prompt cannot spin forever.- Every file path is resolved against the repository root and rejected if it escapes that boundary.
That last point matters more than it sounds. Once you let an agent write to disk in CI, path safety is not an implementation detail. The code explicitly strips line-number suffixes from findings, resolves absolute paths, and rejects anything outside the checked-out repository.
The guardrail I would not remove
What happens after a fix
When a fix is applied, the task creates a branch named from the build ID, commits the changes with a bot identity, pushes the branch, waits until Azure DevOps can see it, and then creates a pull request through the REST API. The PR description includes what was changed, which findings were addressed, and which findings were deliberately skipped for manual review.
security/auto-fix-{BUILD_BUILDID}
git add -A
git commit -m "fix: auto-fix security vulnerabilities"
git push --set-upstream origin security/auto-fix-{BUILD_BUILDID}One design choice I like is that the PR targets the branch that triggered the build, not always main. That keeps the remediation inside the developer's current workflow instead of teleporting changes into a different integration branch.
What the Azure DevOps task exposes
The packaged task surface is intentionally small: endpoint, deployment, optional API key and API version, failure threshold, optional image tag, a flag to skip triage, and a flag to auto-create a dependency-fix PR. Reports land in the artifact staging directory as JSON and Markdown, and findings are also emitted as Azure DevOps log annotations so the failure is visible directly in the pipeline UI.
There is one honest mismatch between the core engine and the packaged task: the Python CLI already contains a richer --auto-fix-iac path, but the task definition I inspected does not expose that flag yet. In other words, the architecture has advanced faster than the extension surface. That is normal for internal platform tooling, but it is worth saying out loud.
The tradeoffs are the whole point
The easy story would be “we used AI to fix security findings”. The real story is about tradeoffs.
- Cost: repeated LLM calls are expensive compared with running one more scanner binary, so the model is used after aggregation, not before.
- Safety: dependency bumps are cheap to validate; arbitrary code changes are not, so the auto-fix scope stays narrow.
- Noise: scanners are deterministic but chatty; the triage pass reduces noise, but it also introduces a probabilistic component that has to be constrained with schema, temperature, and review.
- Coverage: the task detects more ecosystems than it deeply remediates today, which is acceptable as long as that gap is visible.
I think this is the right place for agentic automation in security engineering: not as an oracle, and not as a fully autonomous merger, but as a remediation accelerator sitting behind deterministic scanners and in front of human approval.
What I would change next
The first upgrade is obvious: after a successful fix, I would rerun the relevant scanners before opening the PR. The current loop validates file syntax and structure, which is necessary, but it is not the same thing as proving the original finding is gone. The end-to-end “fix until clean” story is only half true until that rescan is wired in.
The second upgrade is broader scanner coverage. The detector already knows about Node and .NET, but that should turn into first-class scanner paths instead of a future intention. The third is stronger policy around what the agent may edit - for example only dependency manifests, Dockerfiles and explicitly allowlisted IaC directories - because successful automation tends to earn more trust than it should.
But even in its current form, the task proved the thing I cared about: a pipeline can do more than shout. It can classify, explain, propose, and hand a reviewer a concrete patch instead of a red badge.
The stack
- Packaging: Azure DevOps VSIX extension with a Node 20 task entry point.
- Core: Python CLI and shared finding model.
- Scanners: pip-audit, Trivy, TruffleHog, Checkov.
- Triage: Azure OpenAI with structured JSON output.
- Agentic remediation: Microsoft Agent Framework plus validator-backed file edits.
- Integration: Azure DevOps pipeline annotations, artifacts, branch creation and PR automation.