How to Set Up Codex Security After DevDay
Set up Codex Security Cloud, PR review, and the CLI. Current team costs, real Gogs findings, SARIF in CI, and the checks you still need.

You can add repository scans, security reviews on pull requests, and checks before commits without first buying a paid SAST suite, a set of automated security checks against source code. Codex Security now covers those workflows, but the subscription price is only part of the bill: five Standard Business seats cost $125 a month on monthly billing, with scanning usage accounted for separately. Start with one repository and a human who owns its findings.
What You Get After DevDay
Codex Security gives your team three places to catch vulnerable code: the repository, the pull request, and the developer's working copy. Think of them as a building inspection, a check on a proposed renovation, and a check before a worker leaves the site. Each sees a different slice of the same system.
A pull request is a proposed code change awaiting review. The CLI is the command-line tool you run in a terminal; CI is the automated checking process that runs when code changes.
Cloud investigates findings, removes duplicates, and prepares fixes while your computer is closed. OpenAI still calls it a research preview. The September 29 DevDay announcement adds the scheduling and availability context; it does not turn Cloud into a generally available product. DevDay recap, current product status.
GitLab teams should start with the CLI. The current Security Cloud setup connects GitHub. General Codex support for GitLab merge requests does not establish GitLab support for Security Cloud. The separate GitLab CI/CD guide documents a supported way to run security checks on that stack.

Which Plans Get It, and What Five People Pay
Pro, Business, Enterprise, and Edu get Codex Security Cloud, according to the dated DevDay announcement and current Help Center. Security Review lists those same plans and explicitly excludes Plus. Workspace permissions and repository access still have to be enabled. Cloud availability, Security Review access.
For a small company, five Standard Business seats make the relevant subscription comparison:
Token-based billing follows how much the model reads and generates. Seat prices can vary by country and currency. The subscription arithmetic follows the current Business FAQ. Cloud's current billing guidance says existing customers receive notice and must opt in before paid usage begins; scanning pauses without funding. It does not provide an all-in five-person scan quote. Cloud billing, subscription and API pricing distinction.
If you already pay for Business, the added seat cost may be zero. The scan and review usage still belongs in your budget. Your monthly total is the seat bill, plus applicable Cloud usage, API-backed CI scans, runner time, and any optional security dashboard subscription.
Do not budget a new free first month. The free-usage offer was attached to the March 6 launch and covered the next month. No renewed September 30 offer is established by the current product and billing pages. Original launch terms, current billing terms.
For a price anchor, Semgrep's paid Code product lists $30 per contributor per month, or $150 for five contributors. Its Free Edition also supports up to 10 repositories and 10 contributors. A five-person team can therefore evaluate both without pretending every existing scanner requires a purchase. These are different kinds of coverage, so the price comparison alone should not decide your security stack. Semgrep pricing.
Connect One Repo and Get to a Reviewed Fix
Start with an application you understand and a reviewer who can explain its access rules. A threat model is a short description of what the application protects, who can reach it, and where it trusts another system. It is the building plan that tells the inspector which doors should be locked.
- Open Plugins in ChatGPT on desktop or web. Install and enable Codex Security Cloud, then open Security Cloud.
- Select New scan. Connect GitHub if prompted, granting access to the repository you want assessed.
- Choose the repository and a compatible Cloud environment. Create an environment if needed, with the dependencies and test setup the project requires.
- Under What to scan, choose Repository, then Start scan. Follow progress in Scans.
- Open Findings. Read the affected code, validation evidence, and remediation guidance. A validation attempt that failed is unresolved evidence, not a cleared vulnerability.
- Where Fix with Codex is offered, generate the patch, inspect it, and use Create draft pull request after review. Run your normal tests and involve the code owner before merging.
These are the current Cloud setup controls. The environment makes it more practical to reproduce suspected flaws; it does not guarantee that every issue will be validated. Validation behavior.
For ongoing checks, create another scan with Commit changes. Under Repositories, open Monitoring settings to select the environment, set the history window, and enable or pause monitoring. Edit the threat model under Project context as your architecture changes. Cloud also supports scheduled repository scans, as described at DevDay; the current setup walkthrough does not specify schedule cadences or quotas, so there is no daily-scan allowance to assume. Monitoring setup, scheduled scans.
A Real Scanned Repository: Triage the Gogs Findings
Gogs shows why a finding needs deployment context before it becomes a ticket. OpenAI names this repository and the following vulnerabilities among its published Codex Security discoveries. This is a published provider scan case, not a new scan performed for this article. The triage decisions below are recommendations based on the maintainer advisories. OpenAI's disclosed findings.
A CVE number identifies a disclosed vulnerability. Two-factor authentication, or 2FA, adds a second login check beyond a password.
The first advisory lists fixes in 0.13.4 and 0.14.0+dev; the second lists 0.14.0. Those are first-fixed versions from historical advisories, not a recommendation to select an old release today. Check the project's currently supported release before upgrading. Gogs recovery-code advisory, Gogs upload advisory.
Keep the triage record small: deployed revision, reachable entry point, evidence, owner, action, and the test that will verify the fix. A finding can be accepted, dismissed with a concrete reason, or left unresolved for investigation. Only label it fixed after you verify the changed behavior.
Apply the same discipline to your own scan. The CLI's saved reports record findings and coverage. A later scan that omits an earlier finding does not prove it was fixed, and false-positive feedback does not permanently suppress that vulnerability class. Finding history and feedback.
Turn On Automatic Security Review
Make pull-request review your next layer after the initial repository assessment. In Codex settings, choose the repository. Under Review security vulnerabilities, turn on Auto security review and choose All PRs, or personal preferences if you want an opt-in rollout.
Choose On PR open for an initial review, On every push for repeated review as code changes, or Whenever code review runs to pair it with general Code Review. An existing Security Cloud scan is optional. You can reuse its threat model or supply the path to a threat-model file in the repository. Security Review setup.
Automatic reviews report High and Critical findings by default. Manually requested reviews include Medium by default too. You can change those thresholds independently. To request one manually, comment @codex security review on the PR and open the associated task's Security Report for the full evidence. Findings posted to GitHub inherit the PR's visibility.
This is a focused security review. General Code Review can also flag security issues, so expect some overlap. The broader Codex review guide helps you decide where each review belongs.
Install the CLI for Local Work and Pre-Commit
The CLI gives you the same kind of repository investigation in a scriptable package. The source is Apache 2.0 and the public npm package is @openai/codex-security. Public source does not include unrestricted scanning access. Official source and license, CLI prerequisites.
Cloud's included Daybreak Blue model access stays inside Cloud; it does not grant that model access through other Security products or the API. Check your intended sign-in and model access before moving a scan into CI. Product access boundary.
There is a date distinction worth preserving: the GitHub repository was created July 13, 2026; its current public history starts with a July 15 initialization commit, while npm publication starts July 28. July 13 alone is not evidence that the licensed npm CLI shipped that day. Pin the package version when making a repeatable setup. Repository metadata, initial commit, npm publication history.
Use Node.js 22.13.0 or later in the 22 series, Node 24, or Node 26, plus Python 3.10 or later. From your repository, sign in and keep output outside the checkout:
npx @openai/codex-security@0.1.31 login
npx @openai/codex-security@0.1.31 scan . --auth chatgpt \
--output-dir ../codex-security-results --dry-run
npx @openai/codex-security@0.1.31 scan . --auth chatgpt \
--output-dir ../codex-security-resultsReview report.md, findings.json, and coverage.json. Coverage can be complete, partial, or unknown; read deferred areas and open questions even when there are no reported findings. These commands follow the CLI quickstart.
Install the pre-commit check with npx @openai/codex-security@0.1.31 install-hook. It scans staged and unstaged changes, blocks High findings and scan errors by default, and preserves an existing pre-commit script. Because it sees both kinds of changes, keep unrelated experiments out of the working copy when interpreting the result. Hook behavior.
The package version and commands above were checked during this article's preparation. No authenticated local scan, measured scan duration, or scan cost is claimed here.
Put the CLI in CI and Keep the SARIF
SARIF is a standard file format for security findings, so another tool can show each issue beside its source location. Exporting the file and buying a hosted dashboard are separate decisions.
Create a CI secret named CODEX_SECURITY_API_KEY for an account or API organization with the required scan access. The key lets the runner authenticate without an interactive ChatGPT sign-in. This GitHub Actions example scans trusted same-repository PRs, compares their exact base and head, exports SARIF, and preserves the results. It adapts the official CI template, pins the package version verified for this article, and starts with a High-severity gate. Remove --fail-on-severity high for an advisory-only rollout; scan errors and incomplete coverage still need attention.
Save it as .github/workflows/codex-security.yml:
name: Codex Security
on:
pull_request:
jobs:
security:
if: github.event.pull_request.head.repo.full_name == github.repository && github.actor != 'dependabot[bot]'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020
with:
node-version: '26'
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97
with:
python-version: '3.14'
- name: Install trusted CLI outside the checkout
run: npm install --prefix "$RUNNER_TEMP/security-cli" --ignore-scripts --no-audit --no-fund @openai/codex-security@0.1.31
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
with:
ref: ${{ github.event.pull_request.head.sha }}
fetch-depth: 0
persist-credentials: false
- name: Scan and export
env:
OPENAI_API_KEY: ${{ secrets.CODEX_SECURITY_API_KEY }}
CODEX_SECURITY_STATE_DIR: ${{ runner.temp }}/security-state
BASE_SHA: ${{ github.event.pull_request.base.sha }}
HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: |
set -euo pipefail
cli="$RUNNER_TEMP/security-cli/node_modules/.bin/codex-security"
out="$RUNNER_TEMP/security-results"
base="$(git merge-base "$BASE_SHA" "$HEAD_SHA")"
scan_exit=0
"$cli" scan . --diff "$base" --head "$HEAD_SHA" \
--auth api-key --output-dir "$out" \
--fail-on-severity high --json \
> "$RUNNER_TEMP/security-result.json" || scan_exit=$?
if test -f "$out/scan-manifest.json"; then
"$cli" export "$out" --export-format sarif \
--source-root "$GITHUB_WORKSPACE" \
--output "$out/results.sarif"
fi
exit "$scan_exit"
- name: Keep reports, including SARIF when available
if: always()
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
with:
name: codex-security-results
path: |
${{ runner.temp }}/security-results
${{ runner.temp }}/security-result.json
retention-days: 7The workflow preserves the scan exit code while exporting available sealed results. Exit 0 means the selected scope has complete coverage and passes your severity policy. Exit 1 means a finding meets the threshold. Exit 2 means an error or incomplete coverage, including partial or unknown coverage. A green advisory scan is still a report about its selected scope, not a security certificate. Artifact exports and exit codes.

To display SARIF as GitHub code-scanning alerts, add the official template's upload-sarif step and its permissions. Public repositories are supported; private and internal repositories need GitHub Code Security enabled. Keeping SARIF as the artifact above lets you inspect it without promising a free private-repo dashboard. GitHub SARIF eligibility.
For GitLab, use OpenAI's production CI/CD template, which covers merge-request diffs, protected default-branch scans, and opt-in scheduled scans. Its native SARIF ingestion requires GitLab Ultimate 19.2 or later. Ordinary report artifacts are a separate route if that dashboard entitlement is unavailable.
CI runs use the runner's permissions and can inherit its environment. Keep unrelated credentials out of the job and use trusted source changes. Add a --max-cost value once you choose a budget; it is an estimated limit, with possible overshoot from requests already in progress. CI prerequisites, cost controls.
Five Uses Worth Trying, Ranked by Payoff
- A SaaS team changing login or tenant access. Run the baseline, improve the threat model, and enable Security Review on its PRs. The payoff is a better chance of catching an account-boundary mistake before customers encounter it. Keep the auth owner in the review.
- A founder inheriting an older application. Scan the repository once and assign the accepted findings before adding features. This can turn an unknown backlog into a short list of evidenced work, rather than a blanket rewrite request.
- A GitLab team without a managed security scanner. Run CLI checks on merge-request diffs and retain SARIF and coverage as artifacts. The payoff is a repeatable review record within the CI system the team already operates.
- An agency maintaining several client repositories. Use the CLI's resumable bulk scans, separate architecture context, and finding history for each client. This could reduce repeated setup and make a maintenance handover more precise. Bulk scans.
- A team with a noisy existing security backlog. Use Codex Security's backlog-triage workflow to examine scanner results against the current code and controls. The payoff is evidence for which tickets deserve engineering time. Keep the original scanners running. Backlog triage.
Two Services You Could Build Around It
The strongest opportunity is a maintained security rollout for small teams. A founder pays for the setup, threat-model work, CI integration, and recurring human triage, rather than another wrapper around a scan button.
DataForSEO's US Google estimates returned 720 monthly searches for “software vulnerability scanning” and 90 for “code security scanner” in this run. Those are searches, not paying customers, and the broader phrase also includes work outside repository scanning. The $150-a-month five-contributor Semgrep Code price provides a concrete comparison point for a narrowly scoped service. Semgrep price anchor.
The smallest sellable version could cover one owned repository, a documented threat model, the CI workflow, and a reviewed findings queue. Use the customer's access and keep usage costs visible. The catch is operational: you must know enough about the application to reject misleading findings and review sensitive fixes. Selling a security guarantee would overstate the evidence.
A second opportunity is a release evidence pack for agencies. An agency could offer a dated record of scan scope, accepted findings, remaining gaps, and verified remediation for each client handover. The broad query “vulnerability scanning tools” has an estimated 1,900 monthly US searches, but that measures interest in tools across several security domains. It supports a customer-discovery hypothesis, not demand for this exact product.
The MVP could assemble saved scan artifacts and approved triage into a compact client report. The catch is portability and trust: SARIF may travel, but your business value comes from honest interpretation, including incomplete coverage. A generated report alone is easy to copy. All search figures are DataForSEO monthly keyword estimates fetched September 30, 2026; none establishes a growing market or purchase intent.
What It Does Not Replace
Keep dependency scanning, which checks the packages and versions your application imports, and secrets scanning, which detects exposed credentials. Repository reasoning can investigate a related issue, but it is not your complete package inventory or credential-monitoring system.
Keep deterministic SAST where its broad, repeatable checks or your assurance requirements matter. OpenAI explicitly says Codex Security complements SAST. You can try the agent without buying a paid suite; that does not make the existing coverage redundant. Cloud FAQ.
Keep a human on authorization code. Authentication asks who you are; authorization asks which customer's data or action you may access. Those rules depend on business intent and deployment assumptions that a scanner can misunderstand. Have the owner review tenant boundaries, admin permissions, recovery flows, and the regression tests. This is an engineering recommendation, consistent with the vendor's stated need for human threat assessment.
Also keep the scan environment bounded. Cloud uses isolated containers; local and CI scans use your local permissions. The sandbox security guide covers the separate job of controlling what an agent can reach.
Codex or Claude Code's Security Review?
Stay with Claude Code if your immediate need is a security check on pending changes and your team already uses it. Run /security-review locally or configure Anthropic's security-review GitHub Action for PR comments and false-positive filtering. Those features are available to Claude Code users, including paid Pro/Max and API Console accounts. Claude security-review setup.
Choose Codex Security Cloud when the managed repository baseline, commit monitoring, scheduled scans, and prepared fixes are the workflow you want. Choose its CLI when saved finding history, coverage artifacts, and SARIF exports fit your local or CI process. There is no accuracy or cost winner established by this comparison.
Check the filter policy on either stack. Anthropic's security-review Action documents exclusions including denial of service and resource exhaustion, so a disk-exhaustion concern like the Gogs upload case needs deliberate review of that policy. The Action is a separate offering from Claude's hosted Code Review product. Anthropic's security-review repository.
What software is used for vulnerability scanning?
Pick software for the layer you need covered. Codex Security adds repository reasoning and validation. Keep dedicated dependency and secrets checks, and deterministic code scanning where you require that coverage. Network scanning is a different job from reviewing application source.
What is the best free vulnerability scanner?
For this repo-scanning decision, start by separating free source code from free runtime. Codex Security's CLI is Apache 2.0, but scans require access and can consume paid usage. Semgrep also offers a Free Edition within its published repository and contributor limits. Evaluate both against your actual application and review needs. CLI access, Semgrep Free Edition.
Is SonarQube a SAST or DAST?
SonarQube Server is a SAST tool: it examines source code without executing the application. Codex Security's repository reasoning and attempted validation add another kind of investigation alongside established code checks. SonarQube's documented approach.
On Monday, put one owner on one repository. Run the baseline, correct its threat model, triage the first findings, and add advisory CI before choosing a severity gate. Record the usage cost and unresolved coverage before expanding to another repository.
If you want this wired into your team's existing development process, I build production AI systems.
- Last Updated
- Sep 30, 2026
- Category
- Build







