|
| 1 | +--- |
| 2 | +name: dataforseo-python-client |
| 3 | +description: Use the DataForSEO Python client (pip package dataforseo-client, module dataforseo_client) to call DataForSEO API v3 (SERP, Keywords Data, DataForSEO Labs, Backlinks, OnPage, AI Optimization, etc.). Read this before exploring the code; it explains the layout, naming rules and how to find an endpoint without reading the huge generated files. |
| 4 | +--- |
| 5 | + |
| 6 | +# DataForSEO Python client |
| 7 | + |
| 8 | +Generated, typed (pydantic v2) Python client for DataForSEO API v3. |
| 9 | +Every API endpoint is one method; every request/response body is one model class. |
| 10 | + |
| 11 | +- Package: `pip install dataforseo-client`, import as `dataforseo_client` |
| 12 | +- HTTP: `urllib3`, synchronous |
| 13 | +- Base URL: `https://api.dataforseo.com` (sandbox with free dummy data: `https://sandbox.dataforseo.com`) |
| 14 | +- Auth: HTTP Basic with the DataForSEO API login and password (not the dashboard password) |
| 15 | + |
| 16 | +## Do not read generated code in full |
| 17 | + |
| 18 | +The client is generated from an OpenAPI spec and is very large (thousands of model files, `api/*_api.py` files up to ~1 MB, very long field descriptions in every model). Never open files whole. Derive names with the rules below and use targeted search (grep) only to confirm them. |
| 19 | + |
| 20 | +## Layout |
| 21 | + |
| 22 | +Paths are relative to the `dataforseo_client` package (a top-level folder in the repository; after `pip install` it is `site-packages/dataforseo_client/`, where a copy of this file also lives): |
| 23 | + |
| 24 | +``` |
| 25 | +configuration.py Configuration (credentials, host, proxy, ssl, retries) |
| 26 | +api_client.py ApiClient (HTTP transport, context manager) |
| 27 | +exceptions.py ApiException and subclasses |
| 28 | +api/<section>_api.py one class per API section, one method per endpoint |
| 29 | +models/<class_name>.py one pydantic model per file (snake_case file, PascalCase class) |
| 30 | +``` |
| 31 | + |
| 32 | +Sections (the `api/` folder is the source of truth): `SerpApi`, `KeywordsDataApi`, `DataforseoLabsApi`, `DomainAnalyticsApi`, `BacklinksApi`, `OnPageApi`, `ContentAnalysisApi`, `AiOptimizationApi`, `MerchantApi`, `AppDataApi`, `BusinessDataApi`, `AppendixApi`. |
| 33 | + |
| 34 | +## Naming rules (derive names instead of searching) |
| 35 | + |
| 36 | +Endpoint path `/v3/<section>/<rest>` maps to: |
| 37 | + |
| 38 | +| What | Rule | Example for `/v3/serp/google/organic/live/advanced` | |
| 39 | +|---|---|---| |
| 40 | +| API class / module | `<Section>Api` in `api/<section>_api.py` | `SerpApi` in `dataforseo_client.api.serp_api` | |
| 41 | +| Method | snake_case of `<rest>` (usually) | `google_organic_live_advanced` | |
| 42 | +| Request model | `<Section><Rest>RequestInfo` | `SerpGoogleOrganicLiveAdvancedRequestInfo` | |
| 43 | +| Response model | `<Section><Rest>ResponseInfo` | `SerpGoogleOrganicLiveAdvancedResponseInfo` | |
| 44 | +| Task item | `<Section><Rest>TaskInfo` | `SerpGoogleOrganicLiveAdvancedTaskInfo` | |
| 45 | +| Result item | `<Section><Rest>ResultInfo` | `SerpGoogleOrganicLiveAdvancedResultInfo` | |
| 46 | +| Model module | snake_case of class name | `dataforseo_client.models.serp_google_organic_live_advanced_request_info` | |
| 47 | + |
| 48 | +Model names follow the rule strictly. Method names sometimes keep the section prefix (e.g. `dataforseo_labs_id_list`), so confirm the method with one search (the pattern also matches a prefixed name): |
| 49 | + |
| 50 | +```bash |
| 51 | +grep -n "def [a-z_]*google_organic_live_advanced(" api/serp_api.py |
| 52 | +``` |
| 53 | + |
| 54 | +Model fields and their descriptions (required/optional, allowed values, limits) are in `Field(description=...)`. The lines are very long, so list field names first and then grep only the fields you need: |
| 55 | + |
| 56 | +```bash |
| 57 | +grep -oE "^ [a-z_0-9]+:" models/serp_google_organic_live_advanced_request_info.py # field names |
| 58 | +grep -n "^ location_code:" models/serp_google_organic_live_advanced_request_info.py # one field with description |
| 59 | +``` |
| 60 | + |
| 61 | +## Method shapes |
| 62 | + |
| 63 | +- `POST` endpoints: `x(list_optional_x_request_info: List[XRequestInfo]) -> XResponseInfo`, the body is always a list of tasks. Pass it positionally. |
| 64 | +- `GET` endpoints: `x() -> XResponseInfo` or `x(id)` (task id for `task_get_*`, `country` for locations etc.). |
| 65 | +- Every method also has `x_with_http_info(...)` (returns `ApiResponse` with `status_code`, `headers`, `data`) and `x_without_preload_content(...)` (raw urllib3 response). |
| 66 | +- Every method accepts `_request_timeout` (seconds or `(connect, read)` tuple) and `_headers`. |
| 67 | + |
| 68 | +## Setup |
| 69 | + |
| 70 | +```python |
| 71 | +from dataforseo_client import configuration as dfs_config, api_client as dfs_api_provider |
| 72 | +from dataforseo_client.api.serp_api import SerpApi |
| 73 | + |
| 74 | +configuration = dfs_config.Configuration(username="API_LOGIN", password="API_PASSWORD") |
| 75 | +# sandbox: dfs_config.Configuration(host="https://sandbox.dataforseo.com", username=..., password=...) |
| 76 | + |
| 77 | +with dfs_api_provider.ApiClient(configuration) as api_client: |
| 78 | + serp_api = SerpApi(api_client) |
| 79 | + ... |
| 80 | +``` |
| 81 | + |
| 82 | +Create one `ApiClient` and reuse it for all section classes. |
| 83 | + |
| 84 | +## Live request (result in the same call) |
| 85 | + |
| 86 | +```python |
| 87 | +from dataforseo_client import configuration as dfs_config, api_client as dfs_api_provider |
| 88 | +from dataforseo_client.api.serp_api import SerpApi |
| 89 | +from dataforseo_client.rest import ApiException |
| 90 | +from dataforseo_client.models.serp_google_organic_live_advanced_request_info import SerpGoogleOrganicLiveAdvancedRequestInfo |
| 91 | +from pprint import pprint |
| 92 | + |
| 93 | +# Configure HTTP basic authorization: basicAuth |
| 94 | +configuration = dfs_config.Configuration(username='USERNAME',password='PASSWORD') |
| 95 | +with dfs_api_provider.ApiClient(configuration) as api_client: |
| 96 | + # Create an instance of the API class |
| 97 | + serp_api = SerpApi(api_client) |
| 98 | + |
| 99 | + try: |
| 100 | + |
| 101 | + api_response = serp_api.google_organic_live_advanced([SerpGoogleOrganicLiveAdvancedRequestInfo( |
| 102 | + language_name="English", |
| 103 | + location_name="United States", |
| 104 | + keyword="albert einstein" |
| 105 | + )]) |
| 106 | + |
| 107 | + pprint(api_response) |
| 108 | + |
| 109 | + except ApiException as e: |
| 110 | + print("Exception: %s\n" % e) |
| 111 | +``` |
| 112 | + |
| 113 | +## Task-based request (post -> wait -> get) |
| 114 | + |
| 115 | +```python |
| 116 | +from dataforseo_client import configuration as dfs_config, api_client as dfs_api_provider |
| 117 | +from dataforseo_client.api.serp_api import SerpApi |
| 118 | +from dataforseo_client.rest import ApiException |
| 119 | +from dataforseo_client.models.serp_google_organic_task_post_request_info import SerpGoogleOrganicTaskPostRequestInfo |
| 120 | +from pprint import pprint |
| 121 | +import time |
| 122 | + |
| 123 | +# Configure HTTP basic authorization: basicAuth |
| 124 | +configuration = dfs_config.Configuration(username='USERNAME',password='PASSWORD') |
| 125 | + |
| 126 | +def GoogleOrganicTaskReady(id): |
| 127 | + result = serp_api.google_organic_tasks_ready() |
| 128 | + return any(any(xx.id == id for xx in (x.result or [])) for x in (result.tasks or [])) |
| 129 | + |
| 130 | +with dfs_api_provider.ApiClient(configuration) as api_client: |
| 131 | + # Create an instance of the API class |
| 132 | + serp_api = SerpApi(api_client) |
| 133 | + |
| 134 | + try: |
| 135 | + |
| 136 | + task_post = serp_api.google_organic_task_post([SerpGoogleOrganicTaskPostRequestInfo( |
| 137 | + language_name="English", |
| 138 | + location_name="United States", |
| 139 | + keyword="albert einstein" |
| 140 | + )]) |
| 141 | + |
| 142 | + task_id = task_post.tasks[0].id |
| 143 | + |
| 144 | + start_time = time.time() |
| 145 | + |
| 146 | + while GoogleOrganicTaskReady(task_id) is not True and (time.time() - start_time) < 60: |
| 147 | + time.sleep(1) |
| 148 | + |
| 149 | + api_response = serp_api.google_organic_task_get_advanced(id=task_id) |
| 150 | + |
| 151 | + pprint(api_response) |
| 152 | + |
| 153 | + except ApiException as e: |
| 154 | + print("Exception: %s\n" % e) |
| 155 | +``` |
| 156 | + |
| 157 | +Instead of polling you can set `postback_url` / `pingback_url` in the task request. |
| 158 | + |
| 159 | +## Response envelope (same for every endpoint) |
| 160 | + |
| 161 | +``` |
| 162 | +XResponseInfo |
| 163 | + version, status_code, status_message, time, cost, tasks_count, tasks_error |
| 164 | + tasks: List[XTaskInfo] |
| 165 | + XTaskInfo |
| 166 | + id, status_code, status_message, time, cost, result_count, path, data (echo of the request) |
| 167 | + result: List[XResultInfo] # endpoint specific payload, often with items |
| 168 | +``` |
| 169 | + |
| 170 | +- `status_code == 20000` means OK (both top-level and per task); `20100` = task created; `4xxxx`/`5xxxx` = errors. Always check the per-task `status_code`: the HTTP status is usually 200 even when a task failed. |
| 171 | +- All fields are `Optional`; guard lists with `or []`. |
| 172 | +- Models are pydantic: `to_dict()`, `to_json()`, `from_dict()`, `from_json()` are available. |
| 173 | + |
| 174 | +## Polymorphic items |
| 175 | + |
| 176 | +Lists like `items` are typed as a base class (e.g. `BaseSerpApiElementItem`) and deserialized into concrete subclasses by the JSON `type` field (`organic` -> `OrganicSerpElementItem`, `paid` -> `PaidSerpElementItem`, `featured_snippet` -> `FeaturedSnippetSerpElementItem`, ...). Use `isinstance`. The mapping is in `__discriminator_value_class_map` of the base model file; grep there instead of reading it. |
| 177 | + |
| 178 | +## Errors |
| 179 | + |
| 180 | +`dataforseo_client.exceptions` (also re-exported from `dataforseo_client.rest`): `ApiException` with `status`, `reason`, `body`, `headers`; subclasses `BadRequestException` (400), `UnauthorizedException` (401), `ForbiddenException` (403), `NotFoundException` (404), `ServiceException` (5xx). Client-side validation errors raise pydantic `ValidationError`. |
| 181 | + |
| 182 | +## Useful facts |
| 183 | + |
| 184 | +- Location / language codes: `location_code=2840` (United States), `language_code="en"`. Full lists come from endpoints like `SerpApi.google_locations()` / `google_languages()` (and similar per section). |
| 185 | +- Most Live endpoints accept one task per request; Task POST endpoints accept many tasks (up to 100) in one call. |
| 186 | +- `task_get_*` has several variants (`regular`, `advanced`, `html`); use the one matching the data you need. |
| 187 | +- Field semantics, allowed values and limits: the `description` of the model field (it comes from the official API docs). |
| 188 | + |
| 189 | +## External documentation (last resort) |
| 190 | + |
| 191 | +Use https://dataforseo.com/llms.txt only when this file or generated code do not answer the question (for example pricing, account limits or endpoint behaviour that is not described locally). Everything needed to write client code is already in this library. |
| 192 | + |
| 193 | +`llms.txt` is a large (~200 KB) index of links to per-endpoint Markdown pages (`https://docs.dataforseo.com/v3/...md`). Do not read it whole: search it for the endpoint path or name and fetch only the linked page. |
0 commit comments