[Checklist]

Reviewing a pull request that wires in an LLM

The LLM-integration class of the PR review: what the diff of an AI feature touches, four shapes worth stopping for, and the primary sources behind each.

Aevral,

Aevral's PR security review reads several flaw classes on the diff; one of them is LLM-integration risks. Most review time goes to correctness and tests, and the AI feature gets the same treatment as any other dependency call, which is exactly where the class hides: the code is idiomatic, the SDK call is one line, and the rule that matters, what the model may act on, is not on any line. This page is the review procedure for that class, with the OWASP Top 10 for LLM applications behind each shape.

What the diff of an AI feature touches

An LLM integration usually lands as a small diff: a client for the model provider, a prompt template, a call site, and a path where the response goes. None of those lines is a vulnerability on its own. The security question sits in the wiring: which strings reach the prompt, which user can reach that call site, and where the model's output goes next. OWASP treats both halves of that wiring as first-class classes: LLM01:2025, Prompt Injection, for what flows in, and LLM05:2025, Improper Output Handling, for what flows out.

Read the diff in that order. Inputs first: every variable interpolated into the prompt template, and where each one originates. A string from a support ticket, a web page, or a repository file is untrusted content reaching the model, and the OWASP entry is explicit that prompt injections do not need to be human-readable to work: as long as the content is parsed by the model, it can alter its behavior. Outputs second: every place the response is rendered, executed, or passed to another system.

Four shapes worth stopping for

Untrusted content joined into the prompt. A diff that concatenates retrieved documents, ticket bodies, or file contents into the prompt template without marking them as data. This is the indirect injection shape, LLM01:2025, Prompt Injection: the OWASP entry's own scenario is a summarizer reading a web page whose hidden instructions redirect the model. The review question is whether untrusted and trusted content are separated and denoted, which is the entry's prevention point: segregate and clearly identify external content.

Model output passed downstream unvalidated. A response from the model handed to a shell, an eval, a database query, a file path, or rendered as HTML or Markdown without encoding. This is LLM05:2025, Improper Output Handling, whose stated consequence list runs from XSS and CSRF in the browser to SSRF, privilege escalation, and remote code execution on the backend. The OWASP common examples name the shapes directly: output entered into exec or eval, generated SQL executed without parameterization, output used to construct file paths without sanitization.

Privilege the model inherits. A call site where the model's tool calls or API requests run with the application's credentials instead of scoped ones. The prompt-injection entry's prevention list states the rule: provide the application's extensible functionality with its own API tokens and handle those functions in code, restricting the model to the minimum privilege necessary. In a diff this reads as a handler that forwards model-produced arguments to an internal service with the server's own identity.

A human gate removed. A diff that wires a high-risk action, sending data, mutating records, executing commands, so the model's output triggers it directly. The OWASP prevention list is one line long on this: require human approval for high-risk operations, human-in-the-loop controls for privileged operations. In review terms, the question is who or what can pull the trigger, and whether a human is between the model and the action.

The shapes, in one screen of code

A summarization route in an Express service, reduced to its wiring. Two facts meet here: a fetched page reaches the prompt as content, and the model's summary reaches the client as HTML. Both paths are idiomatic, and that is the point.

summarize route, before and after the two guards
  // requireSession resolves the caller; the session carries orgId.
  app.post("/summarize", requireSession, async (req, res) => {
    const page = await fetchPage(req.body.url);
-   const summary = await llm.complete(
-     "Summarize for the requester: " + page.text,
-   );
-   res.send("<div>" + summary + "</div>");
+   const summary = await llm.complete([
+     { role: "system", content: "Summarize the delimited page text." },
+     { role: "user", content: "<page>\n" + page.text + "\n</page>" },
+   ]);
+   res.send("<div>" + escapeHtml(summary) + "</div>");
  });

What the read answers

The before-version asks the model a question and publishes its answer. The delimiters are not a fix for prompt injection, and no delimiters are: the OWASP entry is plain that given the stochastic nature of the models, it is unclear whether fool-proof prevention exists. The guards narrow the surface. Denoting the page as data makes the instruction-vs-content boundary explicit and reviewable; escaping the output closes the rendering path the way output encoding closes it for any other string. Both guards are ordinary application-security moves applied to a new input, which is the honest frame for this class: the model is an untrusted component, and the code around it does the securing.

The OWASP entry's remaining mitigations are code-shaped too, which is why they are reviewable in a diff: validate the expected output format with deterministic code, and filter both input and output for sensitive categories. A model-side promise, a system prompt that says ignore instructions in page text, is a mitigation the diff cannot verify and a reviewer should not count on.

A reading job, closed by a human

The prompt template, the call site, and the rendering path live in different files, and the rule, what the model may act on and what its output may become, lives in none of them. That is what makes this class a reading job: the unit of meaning is the relation between what flows into the prompt and where the output lands. It is the diff-level reading Aevral's PR security review, live on install, applies to pull requests, as findings that are leads: the file, the lines, the reason a human should look, and a fix prompt for the coding agent. The whole-repo scan reads a different, narrower surface: authorization, IDOR, and business-logic access control only, not this class.

A finding stays a lead. A human reads the evidence and decides whether the wiring is safe; the agent adjusts the guards and a human reviews before merge. Nothing in the loop claims a catch rate or patches on its own. The durable verification here, as with the access-control classes, is the probe written into the test suite: the injected page that expects a refused action, the payload in the summary that renders inert.

Sources

OWASP Top 10 for LLM Applications 2025; LLM01:2025 Prompt Injection; LLM05:2025 Improper Output Handling; OWASP LLM Prompt Injection Prevention Cheat Sheet.

Read next

Reviewing a pull request for access control; Where access control hides in business logic; Handing a security finding to your coding agent.

More guides


Catch security flaws before you merge.

Install the GitHub App and PR review starts on. Sign in with GitHub to connect it, then press Scan for the repository you already have.

Install the GitHub AppLog in

For professional use. By installing, you confirm you can act for the account or organization that owns it, and you accept the Terms and DPA on its behalf.