No description
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-10-08 14:23:47 -04:00
.env-example first commit 2026-10-08 14:23:47 -04:00
.gitignore first commit 2026-10-08 14:23:47 -04:00
brave.py first commit 2026-10-08 14:23:47 -04:00
LICENSE first commit 2026-10-08 14:23:47 -04:00
README.md first commit 2026-10-08 14:23:47 -04:00

Brave Search Fetcher

A small Python CLI wrapper around the official Brave Search API. It turns a search query into clean, readable output (or raw JSON) straight from the terminal, talking to the real, sanctioned endpoint: no scraping, no CAPTCHAs, no browser in the middle, and an independent index that is not Google or Bing in disguise.

Designed to be used both interactively by a human and programmatically by scripts, cron jobs, and agents. It is also one half of a metasearch setup: metasearch.py (in the sibling Meta_fetcher project) runs this script and the SearXNG fetcher in parallel and fuses the two result lists via Reciprocal Rank Fusion.


How It Works

Brave Search operates a proper public API with a subscription token. This script calls it directly: one HTTP GET per query, structured JSON back, formatted for humans or piped to jq. There is nothing to scrape and nothing to evade, which is the difference between this wrapper and the SearXNG one: here the search engine is a first-party API, there it is a self-hosted aggregator scraping upstream engines on your behalf.

The flow, end to end:

brave.py "my query"
        │
        ▼
┌─────────────────────┐     loads config from      ┌──────────────────┐
│  brave.py           │ ◄───────────────────────── │  .env            │
│  (CLI front-end)    │                            │  BRAVE_API_KEY   │
└─────────┬───────────┘                            └──────────────────┘
          │ GET api.search.brave.com/res/v1/{web|news|images|videos}/search
          │     X-Subscription-Token: BRAVE_API_KEY
          ▼
┌─────────────────────┐
│  Brave Search API   │  → JSON response
│  (official endpoint)│    (web.results, news.results, ...)
└─────────┬───────────┘
          │
          ▼
┌─────────────────────┐
│  brave.py           │
│  format_results()   │  → human-readable output
│  or --json          │  → raw JSON dump
└─────────────────────┘

Four endpoints exist, one per category: web, news, images, and videos. The script picks the right one from the -c flag, so a category switch is a different API path, not a client-side filter.


The .env File

Configuration lives in a .env file in the same directory as the script. It holds one value:

Variable Purpose
BRAVE_API_KEY Subscription token from the Brave Search API dashboard

A template is provided in .env-example. Copy it to .env and fill in your key:

cp .env-example .env
nvim .env

Example of a filled-in .env file:

BRAVE_API_KEY=BSA_your_token_here

Get the token from the Brave Search API dashboard after subscribing to a plan. The free tier allows 2,000 queries per month at 1 query per second, which is plenty for personal and agent use.

Note: the API key grants whoever holds it your query quota. Treat the .env file as a secret — never commit it. The .gitignore already excludes it.


Requirements

  • Python 3 with requests and python-dotenv:
pip install requests python-dotenv
  • A Brave Search API subscription and its key (see above).

Usage

Basic search:

python3 brave.py "hanshin tigers schedule"

Output example:

1. Hanshin Tigers Official Site
   URL: https://www.hanshintigers.jp/
   The official website of the Hanshin Tigers...
   ⏱ 3 days ago

2. Hanshin Tigers Calendar (August 2026) | NPB.jp
   URL: https://www.npb.jp/...
   ...

Options

Flag What it does
-n, --count N Limit number of results (max 20 per request; default 10)
-c, --category C Search category: web, news, images, videos (default web)
-o, --offset N Pagination offset (max 9; each page up to 20 results)
--country CC 2-letter country code for regional results (e.g. ca, us, jp)
--freshness F Time filter: pd (past day), pw (week), pm (month), py (year)
-j, --json Dump the raw JSON response instead of pretty output
-b, --brief Title + URL only (no snippets) — compact mode

More examples

News category, Canada, past week:

python3 brave.py "quebec general election" -c news --country ca --freshness pw

Images, raw JSON piped to jq:

python3 brave.py "kropotkin portrait" -c images -j | jq '.images.results[].url'

Second page of web results (results 21-40):

python3 brave.py "arch linux kernel 6.18" -n 20 -o 1

Compact output for a quick lookup:

python3 brave.py "mullvad socks5 port" -b

Using it as a library

The module can be imported directly — search() returns the parsed JSON dict:

import sys
sys.path.insert(0, "/path/to/Brave_fetcher")
from brave import search

data = search("awesome wm widgets", count=3)
for r in data["web"]["results"]:
    print(r["url"])

Errors come back as {"error": "..."} rather than exceptions, so callers can degrade gracefully.


How the Code Works

The script is a single file, brave.py, roughly 150 lines, split into four parts:

1. Configuration loading (ENV_PATH, load_dotenv)

ENV_PATH = Path(__file__).parent / ".env"
load_dotenv(ENV_PATH)
API_KEY = os.getenv("BRAVE_API_KEY")
if not API_KEY:
    print("Error: BRAVE_API_KEY not configured in .env file", file=sys.stderr)
    sys.exit(1)

On startup it loads .env from the script's own directory (not the current working directory, so you can call it from anywhere), reads the subscription token, and refuses to run without it — a missing key fails immediately with a clear message instead of a mysterious HTTP 401 later.

2. search() — the API client

def search(query, count=10, category="web", offset=0, country=None, freshness=None) -> dict:

This is the core function. It issues one requests.get() to the category's endpoint with:

  • q — the query
  • count — results per request, clamped to 1-20 (the API's hard limit)
  • offset — pagination, max 9 (up to 10 pages)
  • country / freshness — optional regional and time filters

The key rides in the X-Subscription-Token header. Rate-limit responses (HTTP 429) are detected specifically and returned with the server's Retry-After value surfaced, so callers know how long to wait instead of guessing. Any other network or HTTP error is caught by requests.RequestException and returned as {"error": "..."} — same graceful-error contract as the library example above.

3. format_results() — human output

Turns the JSON dict into the numbered list you see in the terminal:

  • Web results: title, URL, a 200-char snippet, and the result age (⏱) when Brave reports one
  • News results: title plus the source site and age in parentheses (meta_url.source), then URL and snippet
  • Images / videos: title + URL (snippets rarely exist for these)
  • brief=True mode (the -b flag) strips everything to N. Title + URL
  • Errors print as Error: ...; anything unrecognised prints as raw JSON

4. main() — CLI argument parsing

Standard argparse setup mapping the flags above to search() parameters. The -c flag's choices come straight from the CATEGORY_ENDPOINTS map, so the CLI and the API surface can never drift apart.


Troubleshooting

  • Error: BRAVE_API_KEY not configured in .env file — .env missing or empty in the script's directory.
  • 429 Too Many Requests — rate limit hit. The free plan allows 1 query per second; the error message includes the server's Retry-After value. Space your calls or catch and wait.
  • HTTP 401/403 — key invalid, deleted, or the subscription lapsed. Check the API dashboard.
  • Empty output with valid JSON — some queries legitimately return no web block (dictionary-style or navigational queries). Try --json to see the raw response, or a different category.
  • offset beyond 9 — the API caps pagination at 10 pages; refine the query instead of paging deeper.

License

Licensed under the GNU Affero General Public License v3.0 — see LICENSE.

Why AGPL over plain GPL: it's the same strong copyleft, plus it closes the "SaaS loophole" — anyone who modifies this code and offers it as a network service must share their changes, even if they never distribute a binary. For a tool built around self-hosting and open infrastructure, that's the license that actually matches the philosophy. The sibling SearXNG fetcher uses the same license, so the whole search stack speaks the same legal dialect. 🐧