Command injection in agent-written code
Agents work in a terminal all day, so when a feature needs a command-line tool, the first draft is often the exact string they would type, handed to a shell.
Aevral,
In plain words
Imagine leaving a note for a house-sitter: "water the plant named ___". A friend fills in the blank with "fern, then give the spare key to whoever knocks". A house-sitter who follows notes literally does both. Command injection is that note: the application builds a command for the operating system out of text a user supplied, and the shell that runs it treats a semicolon, a pipe or a $( ) in that text as the start of another instruction.
The shell is the problem, not the tool being called. Running the same program with its arguments passed as a list, with no shell in between, leaves those characters as plain text.
Why agents write it: the terminal is their native language
A coding agent spends its session typing shell commands. When a feature needs a thumbnail from ffmpeg, a diff from git or a conversion from pdftotext, the most natural first draft is the command it would type itself, wrapped in exec() or subprocess.run(..., shell=True) with the user's value dropped into a template string. It runs, the output is right, and the test passes with a tidy input.
The same habit shows up in CI. Asked to post a pull request's title in a workflow step, an agent writes the title expression straight into a run: line. The runner substitutes it into the script before the shell sees it, so a pull request titled with a $( ) runs code on the runner.
The shapes it takes in a pull request
A shell string with a user value in it
exec(), execSync(), os.system() or subprocess with shell=True, given a template literal or f-string that includes a filename, timestamp, URL or id from the request.
Hand-made quoting
The value wrapped in quotes to make it safe, which a quote character inside the value undoes.
An argument that becomes an option
No shell, but a user value that starts with a dash is read by the tool as a flag. This is argument injection, a close relative, and a -- separator or a prefix check closes it.
Untrusted text pasted into a CI step
A workflow run: step that interpolates a pull request title, branch name or issue body directly into the script, instead of passing it through an environment variable.
Worked example, an illustration
A video thumbnail endpoint that hands a timestamp to the shell.
The request: let users choose the frame for their video thumbnail. The agent reached for ffmpeg, wrote the command it would run in a terminal, and passed it to the shell. The video lookup is scoped to the owner, so the access check is fine; the timestamp is not.
@app.post("/videos/<int:video_id>/thumbnail")@login_requireddef make_thumbnail(video_id): video = Video.query.filter_by(id=video_id, owner_id=current_user.id).first_or_404() thumb = default_thumbnail(video) ts = request.json.get("timestamp", "00:00:01") out = f"/tmp/thumb-{video.id}.jpg" subprocess.run( f"ffmpeg -y -ss {ts} -i {video.path} -frames:v 1 {out}", shell=True, check=True, ) thumb = out return send_file(thumb)What an Aevral finding on it looks like
On a reviewed pull request, a finding is an inline comment pinned to the added line, next to an advisory Check that never blocks the pull request. This is an illustration of that shape for the diff above.
app/videos.py, added line: `f"ffmpeg -y -ss {ts} -i {video.path} -frames:v 1 {out}"` with shell=True
The timestamp from the request body is formatted into a command string that runs through the shell. A semicolon or $( ) in the value starts a second command on the server.
Missing control: Passing arguments as a list with no shell, and validating the timestamp against the format ffmpeg expects.
Fix with your agent
In app/videos.py, make_thumbnail passes an f-string to subprocess.run with shell=True, and ts comes from request.json. Switch to an argument list without shell=True, and reject ts unless it matches HH:MM:SS with optional milliseconds (return 400). Add a test that posts timestamp="1; touch /tmp/pwned #" and expects a 400 and no file.
Check this finding against the code before changing anything. Make the smallest change that restores the control, add a test for it, and stop before commit or push.A finding is a lead with evidence, not a confirmation. You or your agent check it against the code; Aevral never applies or merges a change. See handing a finding to your coding agent.
How to fix it
Two changes, each doing a different job. Passing the command as a list removes the shell, so a semicolon is just a character in an argument. Checking the timestamp against the one format ffmpeg needs stops odd values from reaching the tool at all, including ones that start with a dash and could be read as an option.
TIMESTAMP = re.compile(r"\d{2}:\d{2}:\d{2}(\.\d{1,3})?")
@app.post("/videos/<int:video_id>/thumbnail")
@login_required
def make_thumbnail(video_id):
video = Video.query.filter_by(id=video_id, owner_id=current_user.id).first_or_404()
ts = request.json.get("timestamp", "00:00:01")
if not isinstance(ts, str) or not TIMESTAMP.fullmatch(ts):
abort(400)
out = f"/tmp/thumb-{video.id}.jpg"
subprocess.run(
["ffmpeg", "-y", "-ss", ts, "-i", video.path, "-frames:v", "1", out],
check=True,
)
return send_file(out)What this page does not claim
- Aevral's PR security review, live on install, looks for command injection on the diff of the pull requests it reviews, as one class among several. Looking for a class does not mean finding each instance of it: a review can miss one, and a quiet review is not proof that the code is clean.
- The whole-repo scan reads authorization, IDOR, and business-logic access control only. It does not read command injection; on this class, the pull request is where Aevral looks.
- No class-specific performance evidence for command injection is published. The receipts page reports named runs of the whole PR review eval suite, not a result for this class.
- The diff, the finding and the fix on this page are illustrations written for this page. They are not output from a customer repository.
- This page covers operating-system commands built by application code and CI steps. Code injection through eval() or a template engine is a different class with different fixes.
Sources
CWE-78: Improper Neutralization of Special Elements used in an OS Command ('OS Command Injection'); CWE-88: Improper Neutralization of Argument Delimiters in a Command ('Argument Injection'); OWASP OS Command Injection Defense Cheat Sheet; Python documentation: subprocess; GitHub Docs: script injections.
Read next
SQL injection in agent-written code; Path traversal in agent-written code; Handing a security finding to your coding agent.
Other classes