API Documentation

Everything you need to integrate Sofya into your AI agent.

llms.txt · SKILL.md

Copy as Markdown
Copy as Text

Quick Start

1. Sign up

Visit the dashboard and sign in with GitHub or with an email link. Your API key is on the dashboard.

2. Get credits

Accounts created by signing in with a GitHub account that is at least 1 year old get the free tier: 2,000 credits, reset to 2,000 every 30 days. Email sign-ups start at 0 credits on pay-as-you-go. Need more? Buy credits from the dashboard at $0.0025 each (minimum $10 = 4,000 credits). See Credits & Billing.

3. Search the web

curl -X POST https://sofya.co/v1/search \
  -H "Authorization: Bearer ay_live_..." \
  -H "Content-Type: application/json" \
  -d '{"query": "latest AI news"}'

Authentication

Every API request needs your API key, sent in either the Authorization header or the X-API-Key header.

Authorization: Bearer ay_live_your_key_here
# or
X-API-Key: ay_live_your_key_here

Exceptions: GET /v1/status and GET /v1/updates need no key. The MCP server also supports OAuth sign-in for clients that offer it. If you set an IP allowlist on a key, requests from any other IP get a 401.

MCP Model Context Protocol

Connect Sofya directly to Claude Code, Cursor, or any MCP-compatible client. Your AI agent gets search, fetch, extract, and research tools, no REST calls needed.

Go to your dashboard and click the copy button for your client. The command is pre-filled with your API key.

Claude Code

claude mcp add --transport http sofya https://mcp.sofya.co/mcp \
  --header "Authorization: Bearer ay_live_..."

Cursor · ~/.cursor/mcp.json

{
  "mcpServers": {
    "sofya": {
      "url": "https://mcp.sofya.co/mcp",
      "headers": { "Authorization": "Bearer ay_live_..." }
    }
  }
}

Codex · ~/.codex/config.toml

[mcp_servers.sofya]
url = "https://mcp.sofya.co/mcp"
http_headers = { "Authorization" = "Bearer ay_live_..." }

Windsurf · ~/.codeium/windsurf/mcp_config.json

{
  "mcpServers": {
    "sofya": {
      "serverUrl": "https://mcp.sofya.co/mcp",
      "headers": { "Authorization": "Bearer ay_live_..." }
    }
  }
}

VS Code Copilot · .vscode/mcp.json

{
  "servers": {
    "sofya": {
      "type": "http",
      "url": "https://mcp.sofya.co/mcp",
      "headers": { "Authorization": "Bearer ay_live_..." }
    }
  }
}

Your API key is sent via HTTP header. The AI model never sees it. Clients that support MCP OAuth can connect to https://mcp.sofya.co/mcp without a header and sign in instead.

Available Tools

search

Web search with page content extraction and optional AI answers. 1-2 credits (+10 with AI answer).

fetch

Fetch URLs and documents as clean markdown. 2 credits per URL; failed URLs are free.

extract

Ask a question about one page and get an AI-written answer. 10 credits.

research

Multi-query deep research with AI synthesis. 50 credits.

MCP tools take the same parameters and return the same JSON as the REST endpoints. Errors arrive as a tool result with isError: true and a message instead of an HTTP status code, and are not charged.

Tool Definitions for Claude & GPT

Copy-paste these tool schemas into your Anthropic or OpenAI API calls. Your model gets Sofya's tools without writing any definitions yourself.

Anthropic (Claude)

Pass this array as the tools parameter in your /v1/messages request. When Claude returns a tool_use block, call the matching Sofya REST endpoint and return the result as a tool_result.

[
  {
    "name": "sofya_search",
    "description": "Search the web for current information. Returns extracted page content, not just snippets. Set topic='news' for current events. Set include_answer=true for an AI-synthesized answer (+10 credits). Returns: query, answer, results [{title, url, content, description, fetched, published_date}], search_depth, topic, elapsed_ms, credits_used, credits_remaining, altered_query, relaxed_query (set when the query matched nothing and was retried once with its site: operator, else its quotes, removed - the results answer that looser query).",
    "input_schema": {
      "type": "object",
      "required": ["query"],
      "properties": {
        "query": {"type": "string", "description": "The search query"},
        "search_depth": {"type": "string", "description": "\"snippets\" (1 credit) or \"basic\" (2, default)"},
        "max_results": {"type": "integer", "description": "Maximum number of results, 1-20 (default 10). Out-of-range values are clamped"},
        "include_answer": {"type": "boolean", "description": "Add AI answer synthesized from results (+10 credits)"},
        "topic": {"type": "string", "description": "\"general\" (default) or \"news\""},
        "freshness": {"type": "string", "description": "\"day\", \"week\", \"month\", \"year\", or \"YYYY-MM-DD:YYYY-MM-DD\""},
        "include_domains": {"type": "array", "items": {"type": "string"}, "description": "Only these domains, bare hostnames like github.com (max 10)"},
        "exclude_domains": {"type": "array", "items": {"type": "string"}, "description": "Exclude these domains, bare hostnames like pinterest.com (max 10)"}
      }
    }
  },
  {
    "name": "sofya_fetch",
    "description": "Fetch one or more URLs in parallel and return their content as clean markdown. Supports web pages, PDF, Word, Excel, PowerPoint, OpenDocument, EPUB, RTF, CSV, XML and plain text. 2 credits per URL, max 10 URLs. Failed URLs (success=false) are not charged. Returns: results [{title, url, content, raw_html, published_time, success, error}], credits_used, credits_remaining.",
    "input_schema": {
      "type": "object",
      "required": ["urls"],
      "properties": {
        "urls": {"type": "array", "items": {"type": "string"}, "description": "URLs to fetch (max 10)"},
        "include_raw_html": {"type": "boolean", "description": "Include raw HTML source in response (default false)"}
      }
    }
  },
  {
    "name": "sofya_extract",
    "description": "Fetch a URL and answer a question about it using AI. Use when you need specific facts from a page (pricing, specs, contact info) rather than its full content. Returns a free-text answer, not structured JSON. 10 credits. If the page has under 200 chars of text it returns empty content with usage.low_content=true instead of a fabricated answer (still charged). Returns: content, url, credits_used, credits_remaining, usage (input_tokens, output_tokens, content_chars, low_content).",
    "input_schema": {
      "type": "object",
      "required": ["url", "prompt"],
      "properties": {
        "url": {"type": "string", "description": "The URL to extract from"},
        "prompt": {"type": "string", "description": "What to extract, e.g. \"list all pricing tiers with features\""}
      }
    }
  },
  {
    "name": "sofya_research",
    "description": "Deep research on a topic. Decomposes query into sub-queries, searches and reads multiple sources in parallel, synthesizes a structured report with citations. 50 credits. Returns: query, report, sources [{title, url, fetched}], sub_queries, credits_used, credits_remaining, usage.",
    "input_schema": {
      "type": "object",
      "required": ["query"],
      "properties": {
        "query": {"type": "string", "description": "The research question or topic"},
        "topic": {"type": "string", "description": "\"general\" (default) or \"news\""},
        "freshness": {"type": "string", "description": "\"day\", \"week\", \"month\", \"year\", or \"YYYY-MM-DD:YYYY-MM-DD\""},
        "max_sources": {"type": "integer", "description": "Max sources to use, 5-30 (default 20). Out-of-range values are clamped"}
      }
    }
  }
]

OpenAI (GPT)

Pass this array as the tools parameter in your /chat/completions request. When the model returns tool_calls, call the matching Sofya REST endpoint and return the result as a role: "tool" message.

[
  {
    "type": "function",
    "function": {
      "name": "sofya_search",
      "description": "Search the web for current information. Returns extracted page content, not just snippets. Set topic='news' for current events. Set include_answer=true for an AI-synthesized answer (+10 credits). Returns: query, answer, results [{title, url, content, description, fetched, published_date}], search_depth, topic, elapsed_ms, credits_used, credits_remaining, altered_query, relaxed_query (set when the query matched nothing and was retried once with its site: operator, else its quotes, removed - the results answer that looser query).",
      "parameters": {
        "type": "object",
        "required": ["query"],
        "properties": {
          "query": {"type": "string", "description": "The search query"},
          "search_depth": {"type": "string", "description": "\"snippets\" (1 credit) or \"basic\" (2, default)"},
          "max_results": {"type": "integer", "description": "Maximum number of results, 1-20 (default 10). Out-of-range values are clamped"},
          "include_answer": {"type": "boolean", "description": "Add AI answer synthesized from results (+10 credits)"},
          "topic": {"type": "string", "description": "\"general\" (default) or \"news\""},
          "freshness": {"type": "string", "description": "\"day\", \"week\", \"month\", \"year\", or \"YYYY-MM-DD:YYYY-MM-DD\""},
          "include_domains": {"type": "array", "items": {"type": "string"}, "description": "Only these domains, bare hostnames like github.com (max 10)"},
          "exclude_domains": {"type": "array", "items": {"type": "string"}, "description": "Exclude these domains, bare hostnames like pinterest.com (max 10)"}
        },
        "additionalProperties": false
      }
    }
  },
  {
    "type": "function",
    "function": {
      "name": "sofya_fetch",
      "description": "Fetch one or more URLs in parallel and return their content as clean markdown. Supports web pages, PDF, Word, Excel, PowerPoint, OpenDocument, EPUB, RTF, CSV, XML and plain text. 2 credits per URL, max 10 URLs. Failed URLs (success=false) are not charged. Returns: results [{title, url, content, raw_html, published_time, success, error}], credits_used, credits_remaining.",
      "parameters": {
        "type": "object",
        "required": ["urls"],
        "properties": {
          "urls": {"type": "array", "items": {"type": "string"}, "description": "URLs to fetch (max 10)"},
          "include_raw_html": {"type": "boolean", "description": "Include raw HTML source in response (default false)"}
        },
        "additionalProperties": false
      }
    }
  },
  {
    "type": "function",
    "function": {
      "name": "sofya_extract",
      "description": "Fetch a URL and answer a question about it using AI. Use when you need specific facts from a page (pricing, specs, contact info) rather than its full content. Returns a free-text answer, not structured JSON. 10 credits. If the page has under 200 chars of text it returns empty content with usage.low_content=true instead of a fabricated answer (still charged). Returns: content, url, credits_used, credits_remaining, usage (input_tokens, output_tokens, content_chars, low_content).",
      "parameters": {
        "type": "object",
        "required": ["url", "prompt"],
        "properties": {
          "url": {"type": "string", "description": "The URL to extract from"},
          "prompt": {"type": "string", "description": "What to extract, e.g. \"list all pricing tiers with features\""}
        },
        "additionalProperties": false
      }
    }
  },
  {
    "type": "function",
    "function": {
      "name": "sofya_research",
      "description": "Deep research on a topic. Decomposes query into sub-queries, searches and reads multiple sources in parallel, synthesizes a structured report with citations. 50 credits. Returns: query, report, sources [{title, url, fetched}], sub_queries, credits_used, credits_remaining, usage.",
      "parameters": {
        "type": "object",
        "required": ["query"],
        "properties": {
          "query": {"type": "string", "description": "The research question or topic"},
          "topic": {"type": "string", "description": "\"general\" (default) or \"news\""},
          "freshness": {"type": "string", "description": "\"day\", \"week\", \"month\", \"year\", or \"YYYY-MM-DD:YYYY-MM-DD\""},
          "max_sources": {"type": "integer", "description": "Max sources to use, 5-30 (default 20). Out-of-range values are clamped"}
        },
        "additionalProperties": false
      }
    }
  }
]

Example: wiring it up

When the model calls a tool, map the tool name to the Sofya endpoint and forward the arguments:

# Python - handle tool calls from Claude or GPT
TOOL_TO_ENDPOINT = {
    "sofya_search": "/v1/search",
    "sofya_fetch": "/v1/fetch",
    "sofya_extract": "/v1/extract",
    "sofya_research": "/v1/research",
}

def call_sofya(tool_name: str, args: dict) -> dict:
    import httpx
    resp = httpx.post(
        f"https://sofya.co{TOOL_TO_ENDPOINT[tool_name]}",
        headers={"Authorization": "Bearer ay_live_..."},
        json=args,
        timeout=120,
    )
    return resp.json()

Core Tools

POST /v1/fetch 2 credits per URL

Fetch up to 10 URLs in parallel and return their content as clean markdown. Works on web pages and documents: PDF, Word (DOCX, DOC), Excel (XLSX, XLS), PowerPoint (PPTX, PPT), OpenDocument, EPUB, RTF, CSV, XML and plain text. Costs 2 credits per URL that succeeds; a URL that fails comes back with success: false and is not charged.

Request Body

{
  "urls": ["https://example.com"],
  "include_raw_html": false
}

urls: required, 1-10 URLs. An empty list is a 400; more than 10 is a 422.

include_raw_html: default false. Also return the page's raw HTML.

Response

{
  "results": [
    {
      "title": "Example Page",
      "url": "https://example.com",
      "content": "# Markdown content...",
      "raw_html": null,
      "published_time": null,
      "success": true,
      "error": null
    }
  ],
  "credits_used": 2,
  "credits_remaining": 1998
}

success / error: each URL succeeds or fails on its own. A URL that hasn't finished after 50 seconds gives up and returns success: false with an error message; the other URLs are unaffected.

Size limits: 5 MB for HTML pages, 50 MB for documents.

raw_html: null unless include_raw_html is true, and may still be null for non-HTML content (such as documents) and for some sites.

published_time: YYYY-MM-DD from page metadata when available, else null.

curl -X POST https://sofya.co/v1/fetch \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer ay_live_..." \
  -d '{"urls": ["https://example.com"]}'
import httpx

resp = httpx.post("https://sofya.co/v1/fetch",
    headers={"Authorization": "Bearer ay_live_..."},
    json={"urls": ["https://example.com"]})
print(resp.json())
const resp = await fetch("https://sofya.co/v1/fetch", {
  method: "POST",
  headers: { "Content-Type": "application/json", "Authorization": "Bearer ay_live_..." },
  body: JSON.stringify({ urls: ["https://example.com"] })
});
console.log(await resp.json());
POST /v1/extract 10 credits

Fetch a webpage and answer a question about it using AI. Costs 10 credits. content is a free-text answer (it is not JSON or structured output; ask for a format in your prompt if you need one). Only the first 50,000 characters of the page text are sent to the model. If the page has under 200 characters of text (empty or JavaScript-rendered body), the model is not called and content is returned empty with usage.low_content: true rather than a fabricated answer. That is still a 200 and is charged 10 credits. If the page itself can't be fetched (404, blocked, no DNS, timed out, refused URL), you get a 422 and are not charged.

Request Body

{
  "url": "https://example.com",
  "prompt": "List all pricing tiers with their monthly prices"
}

url: required. The page to read.

prompt: required, max 4,096 characters. What to find or answer.

Response

{
  "content": "Three tiers: Free ($0), Pro ($20/month), Team ($50/month).",
  "url": "https://example.com",
  "credits_used": 10,
  "credits_remaining": 1990,
  "usage": { "input_tokens": 1420, "output_tokens": 24, "content_chars": 4820, "low_content": false }
}

usage.content_chars is the number of characters of page text the model received; usage.low_content is true when the page had under 200 characters of text (the model was skipped, content is empty, and the call is still charged). Gate on these to detect empty or unrenderable pages.

curl -X POST https://sofya.co/v1/extract \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer ay_live_..." \
  -d '{"url": "https://example.com", "prompt": "Summarize this page"}'
import httpx

resp = httpx.post("https://sofya.co/v1/extract",
    headers={"Authorization": "Bearer ay_live_..."},
    json={"url": "https://example.com", "prompt": "Summarize this page"})
print(resp.json())
const resp = await fetch("https://sofya.co/v1/extract", {
  method: "POST",
  headers: { "Content-Type": "application/json", "Authorization": "Bearer ay_live_..." },
  body: JSON.stringify({ url: "https://example.com", prompt: "Summarize this page" })
});
console.log(await resp.json());
POST /v1/research 50 credits

Deep research on any topic. Decomposes your query into sub-queries, searches and reads multiple sources in parallel, then synthesizes a structured report with citations. Costs 50 credits. Research has a 55-second time limit (504, not charged); if no sources can be found the request fails and is not charged.

Request Body

{
  "query": "How do modern LLMs handle long context?",
  "topic": "general",
  "freshness": null,
  "max_sources": 20
}

query: required, max 4,096 characters.

topic: "general" (default) or "news".

freshness: default null. "day", "week", "month", "year", or "YYYY-MM-DD:YYYY-MM-DD".

max_sources: 5-30, default 20. Out-of-range values are clamped.

Response

{
  "query": "How do modern LLMs handle long context?",
  "report": "## Key Findings\n\n- ...",
  "sources": [
    {
      "title": "Scaling Transformer Context Windows",
      "url": "https://arxiv.org/abs/...",
      "fetched": true
    }
  ],
  "sub_queries": [
    "transformer context window scaling techniques",
    "RoPE positional encoding extensions"
  ],
  "credits_used": 50,
  "credits_remaining": 1950,
  "usage": { "input_tokens": 12400, "output_tokens": 1850 }
}
curl -X POST https://sofya.co/v1/research \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer ay_live_..." \
  -d '{"query": "How do modern LLMs handle long context?"}'
import httpx

resp = httpx.post("https://sofya.co/v1/research",
    headers={"Authorization": "Bearer ay_live_..."},
    json={"query": "How do modern LLMs handle long context?"},
    timeout=120)
print(resp.json())
const resp = await fetch("https://sofya.co/v1/research", {
  method: "POST",
  headers: { "Content-Type": "application/json", "Authorization": "Bearer ay_live_..." },
  body: JSON.stringify({ query: "How do modern LLMs handle long context?" })
});
console.log(await resp.json());

Account

GET /v1/auth/me

Get your account info: credit balance, tier and total requests. Free.

Response

{
  "credits": 1994,
  "plan_credits": 1994,
  "purchased_credits": 0,
  "is_free_tier": true,
  "credits_reset_at": "2026-11-04T12:00:00Z",
  "total_requests": 3,
  "api_key": null,
  "last_login_method": "github",
  "email": "user@example.com",
  "email_verified": true,
  "github_username": "octocat",
  "has_billing_portal": false,
  "new_ip_alerts_enabled": true
}

credits: total spendable balance, plan_credits (free-tier credits) plus purchased_credits.

credits_reset_at: when free-tier credits next reset to 2,000 (every 30 days). Purchased credits never reset.

api_key: always null when you call with an API key. It is only filled in for the signed-in dashboard.

GET /v1/auth/transactions

Get your last 50 top-ups and account adjustments, newest first. Per-request charges and refunds are not listed here; use usage for those.

Response

[
  {
    "id": "uuid",
    "type": "credit",
    "amount": 4000,
    "endpoint": "top-up",
    "balance_after": 5994,
    "created_at": "2026-10-08T12:00:00Z"
  }
]
GET /v1/auth/usage

Get your daily usage breakdown. Rows are per day, endpoint and API key, newest day first, so one day can have several rows for the same endpoint if you use more than one key.

Query Parameters

days: 1-90, default 7. Out-of-range values are clamped.

offset_days: default 0. Skip this many recent days, for paging back in time.

api_key_id: optional. Only count requests made with this key (its id from the dashboard).

curl "https://sofya.co/v1/auth/usage?days=30" \
  -H "Authorization: Bearer ay_live_..."

Response

[
  {
    "date": "2026-10-08",
    "endpoint": "/v1/search",
    "request_count": 3,
    "total_credits": 6
  }
]

Credits & Billing

Every call is paid for in credits. One credit is $0.0025.

Call Credits USD
search (snippets) 1 $0.0025
search (basic) 2 $0.005
include_answer +10 +$0.025
fetch 2 per URL $0.005 per URL
extract 10 $0.025
research 50 $0.125

Buying credits: top up from the dashboard in whole dollars, from $10 (4,000 credits) up to $25,000. Purchased credits are not affected by the free-tier reset. There is no API endpoint for buying credits.

Free tier: accounts created by signing in with a GitHub account that is at least 1 year old get 2,000 credits, reset to 2,000 every 30 days. Unused free credits don't roll over. Linking GitHub to an existing email account does not add the free tier, and email (magic link) sign-ups start at 0 credits on pay-as-you-go.

What is charged: only a 2xx response, at exactly its credits_used. See Billing on errors.

Rate Limits

Requests are rate limited to 30 requests per second per API key. REST (/v1/*) and MCP (/mcp) are counted separately, so each gets its own 30 per second.

If you exceed the limit, the API returns 429 Too Many Requests with a Retry-After header giving the whole number of seconds to wait before retrying.

429 Response

HTTP/1.1 429 Too Many Requests
Retry-After: 1

{
  "detail": "Rate limit exceeded. 30 requests per second."
}

Rate-limited requests do not consume credits. Implement exponential backoff or respect the Retry-After header for best results.

Error Codes

Code Status Description
400 Bad Request A request-level problem the schema can't catch, e.g. an empty urls list on fetch
401 Unauthorized Invalid or missing API key, or the request IP is not in the key's IP allowlist (applies to every /v1 endpoint, including /v1/auth/*)
402 Payment Required Insufficient credits
422 Unprocessable Entity Invalid parameters: a missing or wrong-type field, unknown search_depth or topic, bad freshness format (or over 25 characters), domains that aren't bare hostnames, more than 10 urls or domains, a search query over 2,048 characters, or a research query or extract prompt over 4,096. Also extract when the target page couldn't be fetched (404, blocked, no DNS, timed out, refused URL); check the URL, retrying rarely helps.
429 Too Many Requests Rate limited. Check Retry-After header.
502 Bad Gateway Internal error. Retry the request.
503 Service Unavailable Search capacity is temporarily unavailable. Retry shortly.
504 Gateway Timeout The request hit its time limit (60s; 55s for research). Retry, or for research try a narrower query or fewer sources.

Billing on errors: only a 2xx response is charged, at exactly its credits_used. No error response is charged: credits are reserved when the request starts and refunded automatically if it fails or times out. Inside a 200, some parts are refunded too: on fetch, each URL that comes back success: false; on search, the 10 include_answer credits when no answer was produced (answer: null). Research that finds no sources fails and is refunded. Extract with low_content: true is a 200 and is charged. So on any non-2xx, it is safe to retry.

Over MCP, errors arrive as a tool result with isError: true and a message instead of an HTTP status code. The same billing rules apply.

Status & Monitoring

Per-endpoint health: GET /v1/status (no API key) reports the health of search, fetch, extract and research as measured by Sofya itself.

{
  "overall": "operational",
  "endpoints": {
    "search": { "status": "operational", "latency_ms": 1905.9, "last_check": "2026-10-09T14:00:27+00:00" },
    "fetch": { "status": "operational", "latency_ms": 4116.3, "last_check": "2026-10-09T13:59:47+00:00" },
    "extract": { "status": "operational", "latency_ms": 6437.9, "last_check": "2026-10-09T13:59:18+00:00" },
    "research": { "status": "operational", "latency_ms": 16049.8, "last_check": "2026-10-09T13:59:18+00:00" }
  },
  "incidents": []
}

overall: operational, degraded, partial_outage, major_outage, or checking (while checks have not reported yet).

endpoints.*.status: operational, degraded, down, or unknown. latency_ms is a recent typical latency and may be null. An endpoint can also carry latency_variants (e.g. search snippets vs basic).

incidents: recent incidents.

Changelog: GET /v1/updates (no API key) returns the list of user-facing changes as JSON.

Website status: status.sofya.co runs on a separate server and checks the website (sofya.co) every 60 seconds, so it keeps reporting even if Sofya itself is down. It does not probe the individual API endpoints; use /v1/status for those. Its feeds are public, no API key needed.

status.sofya.co JSON feeds

  • /api/v2/summary.json - Atlassian Statuspage-compatible, for tools that already read Statuspage feeds.
  • /api/v2/status.json - Page metadata and overall indicator only. Cheapest poll if you only need up/down.
  • /api/status.json - Richer feed: last latency, 24h and 90d uptime, last error.
$ curl https://status.sofya.co/api/v2/status.json
{
  "page": { "id": "sofya", "name": "Sofya Status", "url": "https://status.sofya.co", "time_zone": "Etc/UTC", "updated_at": "2026-10-09T14:00:30+00:00" },
  "status": {
    "indicator": "none",
    "description": "All Systems Operational"
  }
}

The indicator field follows the Statuspage convention: none (operational), minor (degraded), or major (down).