Quick Start
1. Sign up
Visit the dashboard and sign in with GitHub or with an email link. Your API key is on the dashboard.
2. Get credits
Accounts created by signing in with a GitHub account that is at least 1 year old get the free tier: 2,000 credits, reset to 2,000 every 30 days. Email sign-ups start at 0 credits on pay-as-you-go. Need more? Buy credits from the dashboard at $0.0025 each (minimum $10 = 4,000 credits). See Credits & Billing.
3. Search the web
curl -X POST https://sofya.co/v1/search \
-H "Authorization: Bearer ay_live_..." \
-H "Content-Type: application/json" \
-d '{"query": "latest AI news"}'
Authentication
Every API request needs your API key, sent in either the Authorization header or the X-API-Key header.
Authorization: Bearer ay_live_your_key_here
# or
X-API-Key: ay_live_your_key_here
Exceptions: GET /v1/status and GET /v1/updates need no key. The MCP server also supports OAuth sign-in for clients that offer it. If you set an IP allowlist on a key, requests from any other IP get a 401.
MCP Model Context Protocol
Connect Sofya directly to Claude Code, Cursor, or any MCP-compatible client. Your AI agent gets search, fetch, extract, and research tools, no REST calls needed.
Go to your dashboard and click the copy button for your client. The command is pre-filled with your API key.
Claude Code
claude mcp add --transport http sofya https://mcp.sofya.co/mcp \
--header "Authorization: Bearer ay_live_..."
Cursor · ~/.cursor/mcp.json
{
"mcpServers": {
"sofya": {
"url": "https://mcp.sofya.co/mcp",
"headers": { "Authorization": "Bearer ay_live_..." }
}
}
}
Codex · ~/.codex/config.toml
[mcp_servers.sofya]
url = "https://mcp.sofya.co/mcp"
http_headers = { "Authorization" = "Bearer ay_live_..." }
Windsurf · ~/.codeium/windsurf/mcp_config.json
{
"mcpServers": {
"sofya": {
"serverUrl": "https://mcp.sofya.co/mcp",
"headers": { "Authorization": "Bearer ay_live_..." }
}
}
}
VS Code Copilot · .vscode/mcp.json
{
"servers": {
"sofya": {
"type": "http",
"url": "https://mcp.sofya.co/mcp",
"headers": { "Authorization": "Bearer ay_live_..." }
}
}
}
Your API key is sent via HTTP header. The AI model never sees it. Clients that support MCP OAuth can connect to https://mcp.sofya.co/mcp without a header and sign in instead.
Available Tools
search
Web search with page content extraction and optional AI answers. 1-2 credits (+10 with AI answer).
fetch
Fetch URLs and documents as clean markdown. 2 credits per URL; failed URLs are free.
extract
Ask a question about one page and get an AI-written answer. 10 credits.
research
Multi-query deep research with AI synthesis. 50 credits.
MCP tools take the same parameters and return the same JSON as the REST endpoints. Errors arrive as a tool result with isError: true and a message instead of an HTTP status code, and are not charged.
Tool Definitions for Claude & GPT
Copy-paste these tool schemas into your Anthropic or OpenAI API calls. Your model gets Sofya's tools without writing any definitions yourself.
Anthropic (Claude)
Pass this array as the tools parameter in your /v1/messages request. When Claude returns a tool_use block, call the matching Sofya REST endpoint and return the result as a tool_result.
[
{
"name": "sofya_search",
"description": "Search the web for current information. Returns extracted page content, not just snippets. Set topic='news' for current events. Set include_answer=true for an AI-synthesized answer (+10 credits). Returns: query, answer, results [{title, url, content, description, fetched, published_date}], search_depth, topic, elapsed_ms, credits_used, credits_remaining, altered_query, relaxed_query (set when the query matched nothing and was retried once with its site: operator, else its quotes, removed - the results answer that looser query).",
"input_schema": {
"type": "object",
"required": ["query"],
"properties": {
"query": {"type": "string", "description": "The search query"},
"search_depth": {"type": "string", "description": "\"snippets\" (1 credit) or \"basic\" (2, default)"},
"max_results": {"type": "integer", "description": "Maximum number of results, 1-20 (default 10). Out-of-range values are clamped"},
"include_answer": {"type": "boolean", "description": "Add AI answer synthesized from results (+10 credits)"},
"topic": {"type": "string", "description": "\"general\" (default) or \"news\""},
"freshness": {"type": "string", "description": "\"day\", \"week\", \"month\", \"year\", or \"YYYY-MM-DD:YYYY-MM-DD\""},
"include_domains": {"type": "array", "items": {"type": "string"}, "description": "Only these domains, bare hostnames like github.com (max 10)"},
"exclude_domains": {"type": "array", "items": {"type": "string"}, "description": "Exclude these domains, bare hostnames like pinterest.com (max 10)"}
}
}
},
{
"name": "sofya_fetch",
"description": "Fetch one or more URLs in parallel and return their content as clean markdown. Supports web pages, PDF, Word, Excel, PowerPoint, OpenDocument, EPUB, RTF, CSV, XML and plain text. 2 credits per URL, max 10 URLs. Failed URLs (success=false) are not charged. Returns: results [{title, url, content, raw_html, published_time, success, error}], credits_used, credits_remaining.",
"input_schema": {
"type": "object",
"required": ["urls"],
"properties": {
"urls": {"type": "array", "items": {"type": "string"}, "description": "URLs to fetch (max 10)"},
"include_raw_html": {"type": "boolean", "description": "Include raw HTML source in response (default false)"}
}
}
},
{
"name": "sofya_extract",
"description": "Fetch a URL and answer a question about it using AI. Use when you need specific facts from a page (pricing, specs, contact info) rather than its full content. Returns a free-text answer, not structured JSON. 10 credits. If the page has under 200 chars of text it returns empty content with usage.low_content=true instead of a fabricated answer (still charged). Returns: content, url, credits_used, credits_remaining, usage (input_tokens, output_tokens, content_chars, low_content).",
"input_schema": {
"type": "object",
"required": ["url", "prompt"],
"properties": {
"url": {"type": "string", "description": "The URL to extract from"},
"prompt": {"type": "string", "description": "What to extract, e.g. \"list all pricing tiers with features\""}
}
}
},
{
"name": "sofya_research",
"description": "Deep research on a topic. Decomposes query into sub-queries, searches and reads multiple sources in parallel, synthesizes a structured report with citations. 50 credits. Returns: query, report, sources [{title, url, fetched}], sub_queries, credits_used, credits_remaining, usage.",
"input_schema": {
"type": "object",
"required": ["query"],
"properties": {
"query": {"type": "string", "description": "The research question or topic"},
"topic": {"type": "string", "description": "\"general\" (default) or \"news\""},
"freshness": {"type": "string", "description": "\"day\", \"week\", \"month\", \"year\", or \"YYYY-MM-DD:YYYY-MM-DD\""},
"max_sources": {"type": "integer", "description": "Max sources to use, 5-30 (default 20). Out-of-range values are clamped"}
}
}
}
]
OpenAI (GPT)
Pass this array as the tools parameter in your /chat/completions request. When the model returns tool_calls, call the matching Sofya REST endpoint and return the result as a role: "tool" message.
[
{
"type": "function",
"function": {
"name": "sofya_search",
"description": "Search the web for current information. Returns extracted page content, not just snippets. Set topic='news' for current events. Set include_answer=true for an AI-synthesized answer (+10 credits). Returns: query, answer, results [{title, url, content, description, fetched, published_date}], search_depth, topic, elapsed_ms, credits_used, credits_remaining, altered_query, relaxed_query (set when the query matched nothing and was retried once with its site: operator, else its quotes, removed - the results answer that looser query).",
"parameters": {
"type": "object",
"required": ["query"],
"properties": {
"query": {"type": "string", "description": "The search query"},
"search_depth": {"type": "string", "description": "\"snippets\" (1 credit) or \"basic\" (2, default)"},
"max_results": {"type": "integer", "description": "Maximum number of results, 1-20 (default 10). Out-of-range values are clamped"},
"include_answer": {"type": "boolean", "description": "Add AI answer synthesized from results (+10 credits)"},
"topic": {"type": "string", "description": "\"general\" (default) or \"news\""},
"freshness": {"type": "string", "description": "\"day\", \"week\", \"month\", \"year\", or \"YYYY-MM-DD:YYYY-MM-DD\""},
"include_domains": {"type": "array", "items": {"type": "string"}, "description": "Only these domains, bare hostnames like github.com (max 10)"},
"exclude_domains": {"type": "array", "items": {"type": "string"}, "description": "Exclude these domains, bare hostnames like pinterest.com (max 10)"}
},
"additionalProperties": false
}
}
},
{
"type": "function",
"function": {
"name": "sofya_fetch",
"description": "Fetch one or more URLs in parallel and return their content as clean markdown. Supports web pages, PDF, Word, Excel, PowerPoint, OpenDocument, EPUB, RTF, CSV, XML and plain text. 2 credits per URL, max 10 URLs. Failed URLs (success=false) are not charged. Returns: results [{title, url, content, raw_html, published_time, success, error}], credits_used, credits_remaining.",
"parameters": {
"type": "object",
"required": ["urls"],
"properties": {
"urls": {"type": "array", "items": {"type": "string"}, "description": "URLs to fetch (max 10)"},
"include_raw_html": {"type": "boolean", "description": "Include raw HTML source in response (default false)"}
},
"additionalProperties": false
}
}
},
{
"type": "function",
"function": {
"name": "sofya_extract",
"description": "Fetch a URL and answer a question about it using AI. Use when you need specific facts from a page (pricing, specs, contact info) rather than its full content. Returns a free-text answer, not structured JSON. 10 credits. If the page has under 200 chars of text it returns empty content with usage.low_content=true instead of a fabricated answer (still charged). Returns: content, url, credits_used, credits_remaining, usage (input_tokens, output_tokens, content_chars, low_content).",
"parameters": {
"type": "object",
"required": ["url", "prompt"],
"properties": {
"url": {"type": "string", "description": "The URL to extract from"},
"prompt": {"type": "string", "description": "What to extract, e.g. \"list all pricing tiers with features\""}
},
"additionalProperties": false
}
}
},
{
"type": "function",
"function": {
"name": "sofya_research",
"description": "Deep research on a topic. Decomposes query into sub-queries, searches and reads multiple sources in parallel, synthesizes a structured report with citations. 50 credits. Returns: query, report, sources [{title, url, fetched}], sub_queries, credits_used, credits_remaining, usage.",
"parameters": {
"type": "object",
"required": ["query"],
"properties": {
"query": {"type": "string", "description": "The research question or topic"},
"topic": {"type": "string", "description": "\"general\" (default) or \"news\""},
"freshness": {"type": "string", "description": "\"day\", \"week\", \"month\", \"year\", or \"YYYY-MM-DD:YYYY-MM-DD\""},
"max_sources": {"type": "integer", "description": "Max sources to use, 5-30 (default 20). Out-of-range values are clamped"}
},
"additionalProperties": false
}
}
}
]
Example: wiring it up
When the model calls a tool, map the tool name to the Sofya endpoint and forward the arguments:
# Python - handle tool calls from Claude or GPT
TOOL_TO_ENDPOINT = {
"sofya_search": "/v1/search",
"sofya_fetch": "/v1/fetch",
"sofya_extract": "/v1/extract",
"sofya_research": "/v1/research",
}
def call_sofya(tool_name: str, args: dict) -> dict:
import httpx
resp = httpx.post(
f"https://sofya.co{TOOL_TO_ENDPOINT[tool_name]}",
headers={"Authorization": "Bearer ay_live_..."},
json=args,
timeout=120,
)
return resp.json()
Core Tools
/v1/search
1-2 credits (+10 with answer)
Search the web. Returns page content, not just snippets. Choose a search depth to control the quality/cost tradeoff. Add include_answer to any depth for an AI-synthesized answer (+10 credits). This is a lightweight alternative to the 50-credit research endpoint.
snippets 1 credit
SERP snippets only. Fastest.
basic 2 credits (default)
Fetches pages, returns extracted content (~12,000 chars per result).
Request Body
{
"query": "latest AI news",
"search_depth": "basic",
"max_results": 10,
"include_answer": false,
"include_domains": [],
"exclude_domains": [],
"topic": "general",
"freshness": null
}
| Parameter | Type | Description |
|---|---|---|
| query | string | Required. Max 2,048 characters. |
| search_depth | string | "basic" (default, 2 credits) or "snippets" (1 credit). |
| max_results | integer | Maximum number of results, 1-20, default 10. Out-of-range values are clamped, not rejected. You may get fewer. |
| include_answer | boolean | Default false. Adds an AI answer synthesized from the results (+10 credits, so basic + answer = 12). If no answer can be produced, answer is null and the 10 credits are refunded. |
| include_domains | string[] | Only return results from these domains. Bare hostnames like github.com (no scheme or path), max 10. |
| exclude_domains | string[] | Leave out these domains. Bare hostnames, max 10. |
| topic | string | "general" (default) or "news". |
| freshness | string | Default null. "day", "week", "month", "year", or a range "YYYY-MM-DD:YYYY-MM-DD". |
Response
{
"query": "latest AI news",
"answer": null,
"results": [
{
"title": "...",
"url": "...",
"content": "Extracted page content...",
"description": "SERP snippet",
"fetched": true,
"published_date": "2026-03-08",
"sublinks": [],
"table": {}
}
],
"search_depth": "basic",
"topic": "general",
"elapsed_ms": 4200,
"credits_used": 2,
"credits_remaining": 1998,
"altered_query": null,
"relaxed_query": null
}
query: the query actually run. It can differ from what you sent: it includes any site: operators built from include/exclude_domains, or the looser query after a zero-result retry.
answer: null unless include_answer is true (and still null, refunded, if no answer could be produced).
content: at basic depth, up to ~12,000 characters of page text per result. A page that doesn't load in time falls back to its snippet, with fetched: false. At snippets depth, fetched is always false.
published_date: YYYY-MM-DD from page metadata or the SERP, or null.
sublinks: extra links shown under a result, each {"title", "description", "url"}. table: an object, often empty.
altered_query: set if the search engine auto-corrected your query, else null.
relaxed_query: set when the query matched nothing and was retried once with its site: operator (else its quotes) removed. The results then answer that looser query. Null otherwise.
topic: "general" (default) for web search, or "news" for news-specific search. Use "news" for current events, breaking news, politics, or any time-sensitive query. Returns articles with publication dates.
freshness: "day", "week", "month", "year", or custom range "YYYY-MM-DD:YYYY-MM-DD"
curl -X POST https://sofya.co/v1/search \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ay_live_..." \
-d '{"query": "latest AI news", "search_depth": "basic"}'
import httpx
resp = httpx.post("https://sofya.co/v1/search",
headers={"Authorization": "Bearer ay_live_..."},
json={"query": "latest AI news", "search_depth": "basic"})
print(resp.json())
const resp = await fetch("https://sofya.co/v1/search", {
method: "POST",
headers: { "Content-Type": "application/json", "Authorization": "Bearer ay_live_..." },
body: JSON.stringify({ query: "latest AI news", search_depth: "basic" })
});
console.log(await resp.json());
/v1/fetch
2 credits per URL
Fetch up to 10 URLs in parallel and return their content as clean markdown. Works on web pages and documents: PDF, Word (DOCX, DOC), Excel (XLSX, XLS), PowerPoint (PPTX, PPT), OpenDocument, EPUB, RTF, CSV, XML and plain text. Costs 2 credits per URL that succeeds; a URL that fails comes back with success: false and is not charged.
Request Body
{
"urls": ["https://example.com"],
"include_raw_html": false
}
urls: required, 1-10 URLs. An empty list is a 400; more than 10 is a 422.
include_raw_html: default false. Also return the page's raw HTML.
Response
{
"results": [
{
"title": "Example Page",
"url": "https://example.com",
"content": "# Markdown content...",
"raw_html": null,
"published_time": null,
"success": true,
"error": null
}
],
"credits_used": 2,
"credits_remaining": 1998
}
success / error: each URL succeeds or fails on its own. A URL that hasn't finished after 50 seconds gives up and returns success: false with an error message; the other URLs are unaffected.
Size limits: 5 MB for HTML pages, 50 MB for documents.
raw_html: null unless include_raw_html is true, and may still be null for non-HTML content (such as documents) and for some sites.
published_time: YYYY-MM-DD from page metadata when available, else null.
curl -X POST https://sofya.co/v1/fetch \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ay_live_..." \
-d '{"urls": ["https://example.com"]}'
import httpx
resp = httpx.post("https://sofya.co/v1/fetch",
headers={"Authorization": "Bearer ay_live_..."},
json={"urls": ["https://example.com"]})
print(resp.json())
const resp = await fetch("https://sofya.co/v1/fetch", {
method: "POST",
headers: { "Content-Type": "application/json", "Authorization": "Bearer ay_live_..." },
body: JSON.stringify({ urls: ["https://example.com"] })
});
console.log(await resp.json());
/v1/extract
10 credits
Fetch a webpage and answer a question about it using AI. Costs 10 credits. content is a free-text answer (it is not JSON or structured output; ask for a format in your prompt if you need one). Only the first 50,000 characters of the page text are sent to the model. If the page has under 200 characters of text (empty or JavaScript-rendered body), the model is not called and content is returned empty with usage.low_content: true rather than a fabricated answer. That is still a 200 and is charged 10 credits. If the page itself can't be fetched (404, blocked, no DNS, timed out, refused URL), you get a 422 and are not charged.
Request Body
{
"url": "https://example.com",
"prompt": "List all pricing tiers with their monthly prices"
}
url: required. The page to read.
prompt: required, max 4,096 characters. What to find or answer.
Response
{
"content": "Three tiers: Free ($0), Pro ($20/month), Team ($50/month).",
"url": "https://example.com",
"credits_used": 10,
"credits_remaining": 1990,
"usage": { "input_tokens": 1420, "output_tokens": 24, "content_chars": 4820, "low_content": false }
}
usage.content_chars is the number of characters of page text the model received; usage.low_content is true when the page had under 200 characters of text (the model was skipped, content is empty, and the call is still charged). Gate on these to detect empty or unrenderable pages.
curl -X POST https://sofya.co/v1/extract \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ay_live_..." \
-d '{"url": "https://example.com", "prompt": "Summarize this page"}'
import httpx
resp = httpx.post("https://sofya.co/v1/extract",
headers={"Authorization": "Bearer ay_live_..."},
json={"url": "https://example.com", "prompt": "Summarize this page"})
print(resp.json())
const resp = await fetch("https://sofya.co/v1/extract", {
method: "POST",
headers: { "Content-Type": "application/json", "Authorization": "Bearer ay_live_..." },
body: JSON.stringify({ url: "https://example.com", prompt: "Summarize this page" })
});
console.log(await resp.json());
/v1/research
50 credits
Deep research on any topic. Decomposes your query into sub-queries, searches and reads multiple sources in parallel, then synthesizes a structured report with citations. Costs 50 credits. Research has a 55-second time limit (504, not charged); if no sources can be found the request fails and is not charged.
Request Body
{
"query": "How do modern LLMs handle long context?",
"topic": "general",
"freshness": null,
"max_sources": 20
}
query: required, max 4,096 characters.
topic: "general" (default) or "news".
freshness: default null. "day", "week", "month", "year", or "YYYY-MM-DD:YYYY-MM-DD".
max_sources: 5-30, default 20. Out-of-range values are clamped.
Response
{
"query": "How do modern LLMs handle long context?",
"report": "## Key Findings\n\n- ...",
"sources": [
{
"title": "Scaling Transformer Context Windows",
"url": "https://arxiv.org/abs/...",
"fetched": true
}
],
"sub_queries": [
"transformer context window scaling techniques",
"RoPE positional encoding extensions"
],
"credits_used": 50,
"credits_remaining": 1950,
"usage": { "input_tokens": 12400, "output_tokens": 1850 }
}
curl -X POST https://sofya.co/v1/research \
-H "Content-Type: application/json" \
-H "Authorization: Bearer ay_live_..." \
-d '{"query": "How do modern LLMs handle long context?"}'
import httpx
resp = httpx.post("https://sofya.co/v1/research",
headers={"Authorization": "Bearer ay_live_..."},
json={"query": "How do modern LLMs handle long context?"},
timeout=120)
print(resp.json())
const resp = await fetch("https://sofya.co/v1/research", {
method: "POST",
headers: { "Content-Type": "application/json", "Authorization": "Bearer ay_live_..." },
body: JSON.stringify({ query: "How do modern LLMs handle long context?" })
});
console.log(await resp.json());
Account
/v1/auth/me
Get your account info: credit balance, tier and total requests. Free.
Response
{
"credits": 1994,
"plan_credits": 1994,
"purchased_credits": 0,
"is_free_tier": true,
"credits_reset_at": "2026-11-04T12:00:00Z",
"total_requests": 3,
"api_key": null,
"last_login_method": "github",
"email": "user@example.com",
"email_verified": true,
"github_username": "octocat",
"has_billing_portal": false,
"new_ip_alerts_enabled": true
}
credits: total spendable balance, plan_credits (free-tier credits) plus purchased_credits.
credits_reset_at: when free-tier credits next reset to 2,000 (every 30 days). Purchased credits never reset.
api_key: always null when you call with an API key. It is only filled in for the signed-in dashboard.
/v1/auth/transactions
Get your last 50 top-ups and account adjustments, newest first. Per-request charges and refunds are not listed here; use usage for those.
Response
[
{
"id": "uuid",
"type": "credit",
"amount": 4000,
"endpoint": "top-up",
"balance_after": 5994,
"created_at": "2026-10-08T12:00:00Z"
}
]
/v1/auth/usage
Get your daily usage breakdown. Rows are per day, endpoint and API key, newest day first, so one day can have several rows for the same endpoint if you use more than one key.
Query Parameters
days: 1-90, default 7. Out-of-range values are clamped.
offset_days: default 0. Skip this many recent days, for paging back in time.
api_key_id: optional. Only count requests made with this key (its id from the dashboard).
curl "https://sofya.co/v1/auth/usage?days=30" \
-H "Authorization: Bearer ay_live_..."
Response
[
{
"date": "2026-10-08",
"endpoint": "/v1/search",
"request_count": 3,
"total_credits": 6
}
]
Credits & Billing
Every call is paid for in credits. One credit is $0.0025.
| Call | Credits | USD |
|---|---|---|
| search (snippets) | 1 | $0.0025 |
| search (basic) | 2 | $0.005 |
| include_answer | +10 | +$0.025 |
| fetch | 2 per URL | $0.005 per URL |
| extract | 10 | $0.025 |
| research | 50 | $0.125 |
Buying credits: top up from the dashboard in whole dollars, from $10 (4,000 credits) up to $25,000. Purchased credits are not affected by the free-tier reset. There is no API endpoint for buying credits.
Free tier: accounts created by signing in with a GitHub account that is at least 1 year old get 2,000 credits, reset to 2,000 every 30 days. Unused free credits don't roll over. Linking GitHub to an existing email account does not add the free tier, and email (magic link) sign-ups start at 0 credits on pay-as-you-go.
What is charged: only a 2xx response, at exactly its credits_used. See Billing on errors.
Rate Limits
Requests are rate limited to 30 requests per second per API key. REST (/v1/*) and MCP (/mcp) are counted separately, so each gets its own 30 per second.
If you exceed the limit, the API returns 429 Too Many Requests with a Retry-After header giving the whole number of seconds to wait before retrying.
429 Response
HTTP/1.1 429 Too Many Requests
Retry-After: 1
{
"detail": "Rate limit exceeded. 30 requests per second."
}
Rate-limited requests do not consume credits. Implement exponential backoff or respect the Retry-After header for best results.
Error Codes
| Code | Status | Description |
|---|---|---|
| 400 | Bad Request | A request-level problem the schema can't catch, e.g. an empty urls list on fetch |
| 401 | Unauthorized | Invalid or missing API key, or the request IP is not in the key's IP allowlist (applies to every /v1 endpoint, including /v1/auth/*) |
| 402 | Payment Required | Insufficient credits |
| 422 | Unprocessable Entity | Invalid parameters: a missing or wrong-type field, unknown search_depth or topic, bad freshness format (or over 25 characters), domains that aren't bare hostnames, more than 10 urls or domains, a search query over 2,048 characters, or a research query or extract prompt over 4,096. Also extract when the target page couldn't be fetched (404, blocked, no DNS, timed out, refused URL); check the URL, retrying rarely helps. |
| 429 | Too Many Requests | Rate limited. Check Retry-After header. |
| 502 | Bad Gateway | Internal error. Retry the request. |
| 503 | Service Unavailable | Search capacity is temporarily unavailable. Retry shortly. |
| 504 | Gateway Timeout | The request hit its time limit (60s; 55s for research). Retry, or for research try a narrower query or fewer sources. |
Billing on errors: only a 2xx response is charged, at exactly its credits_used. No error response is charged: credits are reserved when the request starts and refunded automatically if it fails or times out. Inside a 200, some parts are refunded too: on fetch, each URL that comes back success: false; on search, the 10 include_answer credits when no answer was produced (answer: null). Research that finds no sources fails and is refunded. Extract with low_content: true is a 200 and is charged. So on any non-2xx, it is safe to retry.
Over MCP, errors arrive as a tool result with isError: true and a message instead of an HTTP status code. The same billing rules apply.
Status & Monitoring
Per-endpoint health: GET /v1/status (no API key) reports the health of search, fetch, extract and research as measured by Sofya itself.
{
"overall": "operational",
"endpoints": {
"search": { "status": "operational", "latency_ms": 1905.9, "last_check": "2026-10-09T14:00:27+00:00" },
"fetch": { "status": "operational", "latency_ms": 4116.3, "last_check": "2026-10-09T13:59:47+00:00" },
"extract": { "status": "operational", "latency_ms": 6437.9, "last_check": "2026-10-09T13:59:18+00:00" },
"research": { "status": "operational", "latency_ms": 16049.8, "last_check": "2026-10-09T13:59:18+00:00" }
},
"incidents": []
}
overall: operational, degraded, partial_outage, major_outage, or checking (while checks have not reported yet).
endpoints.*.status: operational, degraded, down, or unknown. latency_ms is a recent typical latency and may be null. An endpoint can also carry latency_variants (e.g. search snippets vs basic).
incidents: recent incidents.
Changelog: GET /v1/updates (no API key) returns the list of user-facing changes as JSON.
Website status: status.sofya.co runs on a separate server and checks the website (sofya.co) every 60 seconds, so it keeps reporting even if Sofya itself is down. It does not probe the individual API endpoints; use /v1/status for those. Its feeds are public, no API key needed.
status.sofya.co JSON feeds
- /api/v2/summary.json - Atlassian Statuspage-compatible, for tools that already read Statuspage feeds.
- /api/v2/status.json - Page metadata and overall indicator only. Cheapest poll if you only need up/down.
- /api/status.json - Richer feed: last latency, 24h and 90d uptime, last error.
$ curl https://status.sofya.co/api/v2/status.json
{
"page": { "id": "sofya", "name": "Sofya Status", "url": "https://status.sofya.co", "time_zone": "Etc/UTC", "updated_at": "2026-10-09T14:00:30+00:00" },
"status": {
"indicator": "none",
"description": "All Systems Operational"
}
}
The indicator field follows the Statuspage convention: none (operational), minor (degraded), or major (down).