Generating SARIF Reports with Sema for CI Integration
Enable Sema's SARIF output to feed static analysis results into GitHub Code Scanning, GitLab SAST, and IDEs. Covers CLI flags, sema.yaml config, a complete GitHub Actions workflow, schema version pitfalls, large-file handling, and a verification checklist.
09 Jul 2025, 23:17 UTC

The Problem: Getting Static Analysis Results Into Your Pipeline
Static analysis tools produce valuable findings, but those findings only matter if they reach developers where they work—in pull requests, IDEs, and CI dashboards. Sema solves this by emitting SARIF (Static Analysis Results Interchange Format), an industry-standard JSON schema that GitHub Code Scanning, GitLab SAST, Azure DevOps, and VS Code all consume natively. This article shows how to enable SARIF output, configure it for your project, and avoid the common pitfalls that break downstream parsers.
Quick Start: One Command
Run Sema with two flags to produce a SARIF file:
sema --report-format sarif --report-file results.sarif
Run this from your repository root (or any directory Sema can analyze). The command writes results.sarif in the current working directory. No elevated permissions are required beyond read access to the source tree.
Expected check: Open results.sarif and verify it starts with {"$schema": "https://schemastore.org/schemas/json/sarif-2.1.0.json", ...} and contains a runs array with at least one results entry.
Persistent Configuration via sema.yaml
For repeatable CI runs, store the settings in a sema.yaml file at the repo root:
report:
format: sarif
file: sema-results.sarif
# Optional: control verbosity
include-passed: false
# Optional: limit to specific rule sets
rule-sets:
- security
- performance
- style
With this file present, sema (no flags) produces sema-results.sarif automatically. The rule-sets list lets you tailor the analysis to your coding standards; omit it to run all built-in rules.
Worked Example: GitHub Actions Integration
The following workflow uploads the SARIF artifact to GitHub Code Scanning. Place it at .github/workflows/sema.yml.
name: Sema Static Analysis
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
sema:
runs-on: ubuntu-latest
permissions:
security-events: write # required for code scanning upload
steps:
- uses: actions/checkout@v4
- name: Install Sema
run: |
curl -fsSL https://sema.dev/install.sh | sh
echo "$HOME/.sema/bin" >> $GITHUB_PATH
- name: Run Sema
run: sema --report-format sarif --report-file sema.sarif
- name: Upload SARIF
uses: github/codeql-action/upload-sarif@v3
with:
sarif_file: sema.sarif
category: sema
Key details: The security-events: write permission is mandatory; without it the upload step fails silently. The category field distinguishes Sema findings from other tools (e.g., CodeQL) in the same repository.
Understanding the SARIF Output
Each finding appears under runs[].results[] with these fields:
ruleId— Sema's internal rule identifier (e.g.,SEMA-SEC-001)level—error,warning, ornotemapped from Sema severitymessage.text— Human-readable descriptionlocations[].physicalLocation— File path, line, and column
Custom rule sets you define in sema.yaml are referenced via tool.driver.rules[], so downstream tools can display your rule names and documentation URLs.
Limits and Common Mistakes
Schema Version Drift
Sema's SARIF schema version is tied to the Sema release. Versions before 1.12 emit SARIF 2.1.0; 1.12+ emit 2.2.0. If your CI parser expects 2.1.0, a newer Sema binary can produce validation errors. Pin the Sema version in your install step (e.g., curl -fsSL https://sema.dev/install.sh | sh -s -- -v 1.11.3) or validate with sarif-validate after every upgrade.
File Size on Large Codebases
A monorepo with 500k lines can generate a 150 MB SARIF file. GitHub Actions has a 100 MB upload limit per artifact. Mitigations:
- Compress before upload:
gzip -c sema.sarif > sema.sarif.gzand upload the .gz (GitHub's uploader accepts gzipped SARIF). - Split by directory using multiple
semainvocations with--pathfilters, then upload each artifact separately.
Color Codes Corrupting JSON
Running sema --no-color --report-format sarif ... is safe, but omitting --no-color when stdout is a TTY can inject ANSI escape sequences into the JSON stream if you redirect incorrectly. Always write directly to --report-file rather than piping stdout.
Duplicate Findings Across Rule Sets
If two rule sets flag the same line (e.g., a security rule and a style rule), SARIF will contain two results entries with identical locations. GitHub Code Scanning deduplicates by ruleId + location, but other tools may show both. Post-process with a small script that merges duplicates by physicalLocation if your dashboard gets noisy.
Verification Checklist
- Local smoke test: Run
sema --report-format sarif --report-file test.sarifon a small sample repo. Opentest.sarifin VS Code (SARIF viewer extension) and confirm rule IDs, severities, and line numbers match the terminal output. - Schema validation: Install
sarif-validate(npm i -g @sarif/validator) and runsarif-validate test.sarif. Exit code 0 means the file conforms to the declared schema. - CI dry run: Push a branch that triggers the workflow above. In the Actions log, verify the "Upload SARIF" step shows "SARIF file uploaded successfully" and the Security tab lists findings under the "sema" category.
- Cross-check with text report: Run
sema --report-format text --report-file text.txtalongside the SARIF run. Compare counts and severities; any discrepancy indicates a conversion bug worth reporting.
When SARIF Isn't Enough
SARIF captures only static, compile-time diagnostics. It does not include runtime profiling data, dependency vulnerability scans, or dynamic taint analysis. If your compliance pipeline requires those, run separate tools and upload their SARIF outputs alongside Sema's—GitHub merges multiple uploads per category automatically.
TL;DR
- Enable SARIF with
--report-format sarif --report-file results.sarifor viasema.yaml. - Upload to GitHub Code Scanning using
github/codeql-action/upload-sarif@v3withsecurity-events: writepermission. - Pin Sema version to avoid schema drift; validate with
sarif-validateafter upgrades. - Compress or split large SARIF files to stay under CI artifact limits.
- Always write to
--report-filedirectly; avoid stdout redirection that can leak ANSI codes.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.