LeakLens is a web-aware secrets scanner for source code, Git history, local files, direct URLs, and modern JavaScript-heavy web applications.
The idea started from the original Praetorian Titus codebase. LeakLens keeps the high-performance secret scanning engine and rule lineage, then extends that foundation into a web-app workflow around crawling, JavaScript discovery, source-map recovery, JS intelligence artifacts, and AI-assisted review.
LeakLens crawling was initially built on ProjectDiscovery Katana and has since evolved into a broader LeakLens-specific discovery and asset-recovery pipeline. The crawler now augments Katana output with initial-response scanning, inline HTML asset extraction, recursive JSON manifest parsing, JavaScript bundle expansion, lazy chunk discovery, same-host URL repair, escaped-path normalization, and downloaded-asset mirroring before scanning. JS intelligence remains a separate informational layer; LeakLens secret findings still come from the scanner rule engine.
Use LeakLens only on codebases, repositories, and websites you are authorized to test.
- CLI scanning for files, directories, Git repositories, direct URLs, and crawled websites.
- Secret detection with optional live validation.
- Katana-based web crawling with LeakLens-specific JS/JSON/source-map discovery and URL repair.
- URL repair for same-host JS paths that are resolved too deeply by crawlers.
- JS intelligence for endpoints, source maps, cloud URLs, subdomains, dependencies, and opt-in dependency-confusion checks.
- Go library usage for embedding the scanner in other internal tools.
LeakLens supports two install paths:
- Install from source with
go install. - Build locally with optional Vectorscan/Hyperscan acceleration.
go install installs from the GitHub repository and writes the binary to $(go env GOPATH)/bin.
Make sure that directory is in your PATH.
Use a tagged release for stable version output:
CGO_ENABLED=1 go install -tags vectorscan github.com/dinosn/leaklens/cmd/leaklens@v0.2.19Use @main to install the latest tested LeakLens branch. The @main examples use GOPROXY=direct so the moving branch is resolved from GitHub instead of a possibly stale Go module proxy response. Main builds display as main@<commit> in leaklens version.
Pure-Go install:
GOPROXY=direct go install github.com/dinosn/leaklens/cmd/leaklens@mainAccelerated install with Vectorscan/Hyperscan:
GOPROXY=direct CGO_ENABLED=1 go install -tags vectorscan github.com/dinosn/leaklens/cmd/leaklens@mainBuild the portable pure-Go binary:
make build-pureThe binary is written to:
dist/leaklensBuild with Vectorscan/Hyperscan acceleration:
make buildThe default make build target uses CGO and the vectorscan build tag. It requires a compatible native library and pkg-config.
Install native dependencies:
# macOS
brew install vectorscan pkg-config
# Ubuntu/Debian
sudo apt-get update
sudo apt-get install -y pkg-config libhyperscan-dev
# Fedora/RHEL
sudo dnf install -y pkgconf-pkg-config hyperscan-develOn macOS, Homebrew installs Vectorscan under the active Homebrew prefix. If pkg-config cannot find libhs, export the package-config path:
export PKG_CONFIG_PATH="$(brew --prefix vectorscan)/lib/pkgconfig:$PKG_CONFIG_PATH"Use matching architectures for Go and the native library. For example, an x86_64 Go toolchain cannot link an arm64 Homebrew Vectorscan library. If the architecture does not match, either install a matching Go toolchain/native library pair or use make build-pure.
On Apple Silicon, this linker error means Go is running as amd64 while Homebrew Vectorscan is arm64:
ld: warning: ignoring file '/opt/homebrew/.../libhs.dylib': found architecture 'arm64', required architecture 'x86_64'
Undefined symbols for architecture x86_64: "_hs_*"
Preferred fix: use an arm64 Go toolchain with the arm64 Homebrew libraries:
brew install go vectorscan pkg-config
export PATH="/opt/homebrew/bin:$PATH"
go env -u GOARCH
go env -u GOOS
export PKG_CONFIG_PATH="$(brew --prefix vectorscan)/lib/pkgconfig:$PKG_CONFIG_PATH"
GOPROXY=direct CGO_ENABLED=1 go install -tags vectorscan github.com/dinosn/leaklens/cmd/leaklens@mainFallback without native acceleration:
GOPROXY=direct CGO_ENABLED=0 go install github.com/dinosn/leaklens/cmd/leaklens@mainCheck the installed build:
leaklens versionCheck GitHub for the latest main branch build:
leaklens updateInstall the latest main branch build directly:
leaklens update --installleaklens update --install preserves the current build mode. A binary built with Vectorscan/Hyperscan runs the vectorscan go install command, while a portable binary runs the normal go install command. Both use GOPROXY=direct so main is resolved from GitHub.
Tagged installs report the tag, such as v0.2.19. Main branch installs report main@<commit> instead of Go's raw pseudo-version.
LeakLens also performs a short main branch check when a command starts. The automatic notification is written to stderr so scan output stays parseable. This matches the documented go install ...@main install path and reports whether the installed binary is built from the latest main commit. If the current build cannot be mapped to a commit, normal scans stay quiet and leaklens update reports that state explicitly.
Disable the automatic check for scripted runs:
leaklens --no-update-check scan path/to/source
LEAKLENS_NO_UPDATE_CHECK=true leaklens scan path/to/source# Scan a file
leaklens scan path/to/file.txt
# Scan a directory
leaklens scan path/to/source
# Scan Git history
leaklens scan --git path/to/repo
# Scan a direct URL
leaklens scan https://example.com/static/app.js
# Scan URLs from a file
leaklens scan --url-file urls.txt
# Crawl a website and scan discovered JS/JSON/source-map files
leaklens scan --crawl https://example.com
# Crawl with JS intelligence
leaklens scan --crawl --js-intel https://example.com
# Validate detected secrets against provider APIs
leaklens scan path/to/source --validateBy default, scan results are printed in human format and kept in memory. Use --output leaklens.ds or another path when you want a datastore for later reporting.
Default web-app crawl profile:
leaklens scan --crawl https://example.comThis default profile enables:
- standard crawling as the primary crawl path
--crawl-browser-capture=truefor a best-effort browser runtime pass when Chrome/Chromium is available--crawl-js-crawl=true--crawl-depth=3--crawl-concurrency=2--crawl-rate-limit=3--crawl-timeout=5m--crawl-max-domain-pages=1000--crawl-extensions=js,json,map--crawl-scope=rdn
Browser runtime capture collects dynamically requested JS/JSON/source-map URLs, bounded text responses, storage snapshots, and common CryptoJS/WebCrypto observations before the normal crawler continues. A plaintext crypto input is reported only when the corresponding operation actually executes and is observable from the page context; values supplied later through a form are not present in a static bundle. Module-local crypto wrappers are still covered by static flow analysis. If Chrome cannot start because the browser is missing, sandboxing blocks it, or an attached DevTools endpoint fails, LeakLens prints one warning and continues with the standard crawler.
The standard crawl always scans the initial response body, including inline JavaScript configuration, even though asset collection defaults to js,json,map. Discovered JSON manifests are inspected recursively for additional matching assets, including ExtJS-style app.json bundles.
Static AES password-flow findings distinguish embedded values from runtime inputs. If a literal password is present in scanned content and passed to a proven password-encryption wrapper, LeakLens reports that literal. If the wrapper receives a form or variable value, LeakLens reports the hard-coded key, mode, padding, wrapper, input expression, and call location without attempting to infer or decrypt a value that is not in the scanned artifacts.
To save crawled files as a readable website-style directory tree while scanning:
leaklens scan --crawl https://example.com/app/ --download-dir downloaded-siteDownloaded URL contents are saved under downloaded-site/<host>/<path>. Query strings are preserved in the filename, so app.js and app.js?v=63 are stored separately.
Use headless crawling only when the standard crawler misses browser-rendered assets:
leaklens scan --crawl --crawl-headless https://example.comFor richer web-app triage:
leaklens scan --crawl --js-intel https://example.comFor applications where JavaScript-discovered relative paths are resolved under the wrong directory, set the application root explicitly:
leaklens scan --crawl --js-intel \
--crawl-base-url https://example.com/app/ \
https://example.com/app/static/js/main.jsLeakLens keeps the crawler-discovered URL as the primary candidate and then tries same-host repaired candidates. When a fallback succeeds, human output shows a repaired: line.
Enable the JS intelligence layer with:
leaklens scan --crawl --js-intel https://example.comThis layer reports supporting artifacts. It does not replace the normal LeakLens secret rules.
| Artifact | Behavior |
|---|---|
| Endpoints | Extracts URL-like values from fetch, importScripts, HTTP method calls, and URL/path properties. |
| Cloud URLs | Finds common S3, GCS, Azure Blob, and DigitalOcean Spaces URLs. |
| Subdomains | Extracts domain-like hostnames from JS/JSON content and tags non-production environment labels such as dev, test, stg, uat, pre, and sandbox. These are naming indicators, not proof that a host is reachable or exposed. |
| Dependencies | Extracts package names from package.json, lockfile package paths, imports, requires, and node_modules references. |
| Source maps | Crawl mode discovers external .map files and probes sibling .js.map files for discovered JavaScript assets. Standalone map files and embedded sourcesContent entries are rescanned with normal LeakLens rules. JS intelligence also decodes inline source maps when enabled. |
| Generic secret heuristic | Optional low-confidence JS assignment heuristic. Values are masked in JS-intel output. |
| Dependency-confusion candidate | Optional active npm registry check. Reports packages that return 404 from the configured registry. |
Useful examples:
# Include masked low-confidence generic JS assignments
leaklens scan --crawl --js-intel --js-intel-generic-secrets https://example.com
# Actively check discovered npm packages for public-registry misses
leaklens scan --crawl --js-intel --js-intel-npm-check https://example.com
# Disable inline source-map parsing while keeping other JS intelligence
leaklens scan --crawl --js-intel --js-intel-source-maps=false https://example.comJS-intelligence artifacts are persisted in the LeakLens datastore independently of secret findings. Direct scan --format json and SARIF output remain secret-only for compatibility. Use leaklens report --include-js-intel for a human report, or combine it with --format json for an opt-in {findings, js_intel} envelope.
leaklens scan [target] [flags]Targets can be:
- A local file or directory.
- A local Git repository.
- A direct HTTP(S) URL.
- A GitHub repository reference such as
github.com/org/repo. - A GitLab project reference such as
gitlab.com/group/project.
| Flag | Default | Description |
|---|---|---|
--output |
:memory: |
Output datastore path. Use a path such as leaklens.ds when you want a persistent datastore. |
--format |
human |
Output format: human, json, or sarif. |
--rules |
Path to a custom rule file or directory. | |
--rules-include |
Include rules matching regex patterns, comma-separated. | |
--rules-exclude |
Exclude rules matching regex patterns, comma-separated. | |
--git |
false |
Treat the target as a Git repository and enumerate history. |
--max-file-size |
20971520 |
Maximum file size to scan. Accepts bytes or KB, MB, GB. |
--include-hidden |
false |
Include hidden files and directories. |
--context-lines |
3 |
Lines of context before and after matches. Use 0 to disable. |
--incremental |
false |
Skip already-scanned blobs. |
--validate |
false |
Validate detected secrets against provider APIs. |
--validate-workers |
4 |
Concurrent validation workers. |
--workers |
CPU count | Parallel scan workers. |
--store-blobs |
false |
Store file contents under the datastore blob directory. |
--download-dir |
Write downloaded URL contents to a directory that preserves the website path structure. | |
--url-file |
File containing URLs to scan, one per line. Use - for stdin. |
|
--sqlite-row-limit |
1000 |
Max rows per SQLite table when extraction is enabled. Use 0 for unlimited. |
| Flag | Default | Description |
|---|---|---|
--extract |
Extract text from binary files. Supported values include xlsx, docx, pdf, zip, or all. |
|
--extract-max-size |
10MB |
Max uncompressed size per extracted file. |
--extract-max-total |
100MB |
Max total bytes to extract from one archive. |
--extract-max-depth |
5 |
Max nested archive depth. |
| Flag | Default | Description |
|---|---|---|
--crawl |
false |
Enable crawl mode. |
--crawl-depth |
3 |
Maximum crawl depth. |
--crawl-concurrency |
2 |
Concurrent crawl workers. |
--crawl-rate-limit |
3 |
Maximum crawl requests per second. |
--crawl-host-rate-limit |
0 |
Maximum requests per second per host. 0 uses --crawl-rate-limit. |
--crawl-timeout |
5m |
Maximum crawl duration. |
--crawl-browser-capture |
true |
Best-effort browser runtime pass for dynamic assets, text responses, storage, and crypto calls that execute and are observable from the page context. Falls back to the standard crawl with one warning if Chrome is unavailable. |
--crawl-headless |
false |
Use a headless browser for JS-heavy sites. |
--crawl-js-crawl |
true |
Parse JavaScript files for additional endpoints. |
--crawl-extensions |
js,json,map |
File extensions to collect and scan. |
--crawl-scope |
rdn |
Scope: rdn, dn, or fqdn. |
--crawl-base-url |
Application base URL for repairing JS-discovered relative paths. | |
--crawl-max-domain-pages |
1000 |
Per-host page and discovery safety cap. A hit emits a probable-fallback-recursion warning; set 0 for unlimited. |
--crawl-chrome-data-dir |
Chrome user-data-dir for preserving sessions. | |
--crawl-chrome-ws-url |
Chrome DevTools websocket URL for attaching to an existing browser. | |
--crawl-system-chrome-path |
Chrome or Chromium binary path. | |
--crawl-use-installed-chrome |
false |
Use installed Chrome instead of Katana-managed Chrome. |
--crawl-no-incognito |
false |
Run headless crawl without an incognito context. |
--crawl-no-sandbox |
false |
Run headless Chrome with --no-sandbox. Auto-enabled when LeakLens launches headless Chrome as root. |
--crawl-automatic-form-fill |
false |
Enable Katana automatic form filling and submission. |
--crawl-auth |
username:password for Katana automatic login. |
| Flag | Default | Description |
|---|---|---|
--js-intel |
false |
Extract JS intelligence artifacts and rescan inline source-map sources. |
--js-intel-source-maps |
true |
Parse inline source maps and scan embedded sources when --js-intel is enabled. |
--js-intel-generic-secrets |
false |
Enable low-confidence JS-style generic secret heuristics. |
--js-intel-npm-check |
false |
Actively check discovered npm packages for public-registry misses. |
LeakLens can run an optional AI-assisted review after the normal deterministic scan. This is intended for authorized validation and remediation work over JavaScript, TypeScript, JSON, and source-map artifacts. The scanner rules still run first and remain the deterministic source of secret findings. The AI layer is a second-stage analyst that reviews the collected files, explains application behavior, proposes additional secret candidates, and writes owner-facing test plans.
AI review works with crawled sites, direct URL scans, URL-file scans, local files, local directories, local Git repositories, and cloned GitHub/GitLab targets. For URL and crawl scans, --ai automatically enables downloaded-file storage. If --download-dir is not supplied, LeakLens creates one under the AI report directory.
Provider configuration is environment-only so API keys are not exposed in command lines:
export LEAKLENS_AI_PROVIDER=openai # openai or anthropic
export LEAKLENS_AI_MODEL=your-model-name
export LEAKLENS_OPENAI_API_KEY=...
export LEAKLENS_ANTHROPIC_API_KEY=...
export LEAKLENS_AI_TIMEOUT=5m # optional per provider request timeout
export LEAKLENS_AI_RETRIES=3 # optional transient provider retry count
export LEAKLENS_AI_CHUNK_CHARS=60000 # optional max redacted characters per AI chunk
export LEAKLENS_AI_CONCURRENCY=3 # optional parallel AI chunk reviewsExamples:
leaklens scan --crawl --ai https://example.com
leaklens scan --ai --ai-mode secrets path/to/downloaded-site
leaklens scan --ai --ai-mode appsec --ai-report-dir reports/example path/to/app
leaklens scan --crawl --ai --ai-cloud-redaction expanded https://example.comAI flags:
| Flag | Default | Description |
|---|---|---|
--ai |
false |
Run AI-assisted JS/JSON/source-map analysis after scanning. |
--ai-mode |
all |
AI analysis mode: secrets, appsec, or all. |
--ai-report-dir |
Directory for AI artifacts. Default: leaklens-ai/<target>-<timestamp>. |
|
--ai-cloud-redaction |
standard |
Cloud redaction mode: standard or expanded. Target URLs and hostnames are always redacted in both modes. |
--ai-progress |
text |
AI progress output: text or quiet. A JSON-lines progress artifact is always written. |
--ai-resume |
false |
Reuse completed AI response checkpoints from the same --ai-report-dir. |
AI artifacts:
| File | Description |
|---|---|
leaklens-ai-report.md |
Markdown report with scope, coverage, AI findings, curl-style validation ideas, locations, and remediation notes. |
corpus-manifest.json |
Local manifest of every file selected for AI review, its cloud-redacted path, size, line count, and chunk count. |
ai-progress.ndjson |
Machine-readable progress events for long-running AI analysis. |
ai-redaction-map.json |
Local-only mapping from redacted cloud placeholders back to real origins, hostnames, and file paths. Do not upload this file to third parties. |
ai-chunks/ |
Checkpointed provider responses for completed overview and file-chunk reviews. These are reused by --ai-resume. |
AI provider resilience:
- Transient provider failures such as request timeouts, HTTP 408, HTTP 429, and HTTP 5xx responses are retried according to
LEAKLENS_AI_RETRIES. - A failed overview or file chunk no longer aborts the whole scan after the local corpus is built. LeakLens writes a partial Markdown report, records failed stages in
AI Failures, and preserves successful responses inai-chunks/. - To continue a long run, rerun the same scan with the same
--ai-report-dir --ai-resume. Completed checkpoints are reused and only missing chunks are sent to the provider. - File chunks are reviewed in parallel according to
LEAKLENS_AI_CONCURRENCY. Set it to1for serial provider calls or lower it if the provider rate-limits the run. - With
--ai-progress=text, high-signal AI observations from completed overview and chunk responses are printed immediately asAI insight:lines while the full report is still being built. - Use
LEAKLENS_AI_TIMEOUTfor slow provider responses andLEAKLENS_AI_CHUNK_CHARSto reduce request size for models or networks that time out on large chunks. The default chunk size is60000characters, matching the original AI chunking behavior.
Cloud redaction behavior:
- Target URLs and hostnames are always redacted before provider submission. This is not bypassable by flags.
- URL paths are preserved so AI can still reason about endpoints. For example,
https://www.example.com/api/admin/configis sent asTARGET_ORIGIN_1/api/admin/config. - Third-party origins are also redacted, for example
https://cdn.vendor.com/lib.jsbecomesEXTERNAL_ORIGIN_1/lib.js. - Local absolute paths are replaced with file placeholders before provider submission.
--ai-cloud-redaction standard redacts obvious credentials, authorization headers, cookie values, secret-bearing query values, and generic high-entropy strings before provider submission. It keeps variable names, function names, endpoint paths, HTTP methods, and structural context.
--ai-cloud-redaction expanded still redacts target origins, hostnames, local filesystem paths, authorization headers, cookie values, and obvious credential assignments, but preserves URL endpoint paths and more non-URL JavaScript context. For example, https://www.example.com/api/files/upload is sent and reported as TARGET_ORIGIN_1/api/files/upload, not as a generic redacted upload endpoint. Use this mode when deeper AI reasoning over config and string literals is needed. It does not disable target origin or hostname redaction.
LeakLens does not execute AI-generated curl commands. The report contains validation plans only. Active testing may be added later behind a separate explicit option.
Direct repository references work through scan:
leaklens scan github.com/org/repo
leaklens scan gitlab.com/group/projectUse dedicated commands for org, user, group, and token-aware workflows:
# GitHub repository
leaklens github org/repo
# GitHub organization
leaklens github --org my-org --token "$GITHUB_TOKEN"
# GitLab project
leaklens gitlab group/project
# GitLab group
leaklens gitlab --group my-group --token "$GITLAB_TOKEN"Important flags:
| Flag | GitHub | GitLab | Description |
|---|---|---|---|
--token |
yes | yes | API token. Optional for public projects. |
--git |
yes | yes | Scan full Git history. |
--no-clone |
yes | yes | Fetch files via API instead of cloning. Requires token and does not scan Git history. |
--org |
yes | no | Scan all repositories in a GitHub organization. |
--user |
yes | yes | Scan repositories/projects for a user. |
--group |
no | yes | Scan all GitLab projects in a group. |
--url |
no | yes | GitLab base URL. Default is gitlab.com. |
--output |
yes | yes | Output database path. Default is leaklens.db. |
--format |
yes | yes | Output format: human or json. |
Scan results are kept in memory unless --output is set to a datastore path.
# Human report from leaklens.ds
leaklens report
# JSON report
leaklens report --format json
# Include persisted JS intelligence in a human report
leaklens report --include-js-intel
# Opt in to the combined machine-readable report envelope
leaklens report --format json --include-js-intel
# SARIF report
leaklens report --format sarif
# Report from a specific datastore
leaklens report --datastore path/to/leaklens.ds
# TUI exploration
leaklens explore --datastore path/to/leaklens.ds# List built-in rules
leaklens rules list
# Scan with only matching rules
leaklens scan path/to/source --rules-include "aws,gcp"
# Exclude noisy rules
leaklens scan path/to/source --rules-exclude "generic"
# Use custom rules
leaklens scan path/to/source --rules path/to/rules.yamlgo get github.com/dinosn/leaklenspackage main
import (
"fmt"
"log"
"github.com/dinosn/leaklens"
)
func main() {
scanner, err := leaklens.NewScanner()
if err != nil {
log.Fatal(err)
}
defer scanner.Close()
matches, err := scanner.ScanString(`aws_access_key_id = AKIAIOSFODNN7EXAMPLE`)
if err != nil {
log.Fatal(err)
}
for _, match := range matches {
fmt.Printf("%s at offset %d\n", match.RuleName, match.Location.Offset.Start)
}
}LeakLens supports:
human: grouped per-file output with a summary.json: machine-readable matches.sarif: SARIF 2.1.0 output with tool nameleaklens.
Examples:
leaklens scan path/to/source --format json
leaklens scan path/to/source --format sarifValidation is disabled by default. When enabled, LeakLens attempts provider-specific checks for supported secret types:
leaklens scan path/to/source --validate --validate-workers 8Validation can make outbound requests to provider APIs. Use it only when that is acceptable for the engagement.
# Pure-Go portable binary
make build-pure
# Accelerated build when compatible Vectorscan/Hyperscan is installed
make build
# Static binary
make build-static
# Tests
go test ./...If macOS blocks the default Go cache in a restricted environment, use explicit caches:
GOCACHE=/private/tmp/go-build-leaklens \
GOMODCACHE=/private/tmp/go-mod-leaklens \
go test ./...LeakLens builds on and credits:
- Praetorian Titus: original scanner foundation and Go implementation lineage.
- NoseyParker and Kingfisher: detection-rule lineage.
- ProjectDiscovery Katana: crawling engine used by
--crawl. - PortSwigger js-miner: inspiration for JS intelligence concepts such as endpoint, dependency, source-map, and JavaScript artifact discovery.
LeakLens is maintained independently in this repository.
This project follows the repository license. Keep upstream attribution when redistributing derivative work.