Devin’s documentation for Security Swarm describes a scanner that sends many AI agents through a codebase at once, and it never says how often those agents are wrong. That absence matters more than any feature on the page. A security team buried in bogus findings stops reading them, so precision decides whether a scanner survives its first quarter. The page offers no false-positive rate, no precision figure, and no benchmark numbers.
This is product documentation from Cognition, the company behind Devin, not an announcement. The page carries no date, no price, and no word on which plans include the feature. Scans consume Devin’s usage units, called ACUs, but the docs give no figure for what a scan costs.
The list of targets is broad. Remote code execution means an attacker getting a server to run their commands. SQL injection means slipping database instructions in through an ordinary input field. Server-side request forgery, or SSRF, means tricking a server into making requests on the attacker’s behalf. The docs add denial of service, which knocks a system offline, along with memory-safety bugs, authorization bypasses, path traversal, and exploits chained across several files. No programming languages are named anywhere, so a buyer cannot tell from the page whether their stack is covered.
Cognition calls the method Agentic MapReduce, its name for splitting a repository across parallel Devin sessions and combining what they find. The docs state the goal, broad coverage with deep investigation at bounded cost, but do not describe how the pieces are divided or merged. The only mechanical hint is a batch-size setting that groups files “with signals” into investigation batches, five by default and adjustable from one to 500. What counts as a signal is not defined. The docs also say Cognition tested Security Swarm against published vulnerabilities from the GitHub Advisory Database, and then report no results.
On false positives, the documentation describes process rather than outcomes. Devin first drafts a threat model for the repository, and in interactive mode it pauses so a person can approve or correct it before any investigation starts. Profiles can instruct Devin to separate vulnerabilities an attacker can actually reach from theoretical ones. Each finding carries severity, exploitability, and confidence ratings, though the docs do not say what the confidence scale means or how it is calibrated.
The strongest check is optional sandbox validation, where a separate Devin session builds the application and tries to demonstrate the flaw, attaching the result and artifacts to the finding. It runs only if the profile enables it and includes instructions for building, seeding, and authenticating the app. By default it covers critical, high, and medium findings. Beyond that, filtering is manual: a status called Dismissed exists for false positives and duplicates, and a feedback action lets a human teach the profile about controls the scan missed.
Operationally, a scan is started from the Security page, a slash command inside a session, an automation, or Cognition’s API. Schedules can rescan only commits added since the last run, and automations accept event triggers such as webhooks. Continuous integration is not mentioned by name. Scans can cover one repository, up to 200 analyzed together, or every matching repository in an organization, each on its own. Nothing is fixed automatically. A person assigns a finding to Devin, which then opens a pull request, and findings can be exported as CSV. Members get no scan permissions unless an administrator grants them.
Cognition sells agents that write code, and now an agent to audit it, so the reviewer and the author may share a lineage. The docs themselves tell buyers to compare Security Swarm with another scanner using identical scope, threat model, and severity criteria. Run that comparison on your own repository, and count how many findings your engineers dismiss.
Devin documentation for Security Swarm, published by Cognition at docs.devin.ai (undated product documentation).