A standalone Node script that repairs DatoCMS records whose localized SEO field had its source locale overwritten by a target language.
Versions <= 3.4.5 of the ai-translations plugin had a bug: translating
a SEO field to more than one target locale at once could overwrite the source
locale with one of the targets (e.g. an en -> es run wrote Spanish into the
en slot). The plugin bug is fixed as of 3.4.6. This tool cleans up records
that were already corrupted, bringing each one back to a last-known-good state
of the source locale and blanking the other locales so they can be
re-translated. Optionally, it can instead use DeepL to backfill every locale
with "good enough" SEO in one pass (see Translation backfill).
It uses the DatoCMS CMA client
to read/write records and version history, and
eld (Efficient Language Detector) to
decide which language a SEO value is actually in.
- Only the SEO field(s) of each record. No other field is ever read for writing or included in an update payload.
- For each record of each configured model it:
- Reads the SEO field's source-locale
title/description. - Detects the language with
eld, restricted to the environment's locales. - If it is reliably the source language, leaves the record alone.
- Otherwise walks the record's version history to find the last
known-good source-locale SEO and reverts only the source locale to it,
while blanking the title/description of the other locales (their
image/twitter_card/no_indexare preserved).
- Reads the SEO field's source-locale
- Records are left as drafts — nothing is published.
Pre-flight safety check. Because step 4 blanks the
title/descriptionof non-source locales, the tool is incompatible with a SEO field that has the "required SEO fields" validation enabled fortitleordescription: the API would reject every revert (VALIDATION_REQUIRED_SEO_FIELDS). The tool detects this from the schema and aborts before reading or writing anything, listing the offending field(s). To proceed, temporarily disable that validation (model settings → the SEO field → Validations), run the tool, then re-enable it.Exception —
AUTOMATICALLY_TRANSLATE. In translate mode the tool fills every locale rather than blanking it (keep / move / revert / DeepL-backfill), so the validator is no longer endangered. The check is therefore downgraded from an abort to a per-field warning and the run proceeds — no need to disable the validation. (One caveat: if a record's only known-good source genuinely lacks a required sub-field, that single write can still be rejected and is reported as a 🚫 for that record; the rest of the run is unaffected.)
- Primary — the most recent version whose source SEO is reliably the source language.
- Fallback — if no version is reliably the source language, the most recent version whose source SEO is populated (and not reliably a foreign language) while every other locale is empty — i.e. the original, human-authored, pre-translation state. Reverts via the fallback are flagged Double-check.
- Move — if neither version exists but the source SEO is reliably a
foreign language whose own locale is currently blank, the content is
simply in the wrong place: its
title/descriptionare moved into that locale and cleared from the source. The reverse never happens — the source locale is never filled from a value detected in another locale. Flagged Double-check. - If none of the above apply, the record is left untouched and flagged for manual review.
To avoid clobbering legitimate short titles, records whose source value is only
low-confidence source language (too short for eld to be sure, but its best
guess is the source language) are left untouched.
A schema pre-flight runs first (see the note above); if it passes, every
record is taken through the tree below. Only the four WRITE leaves change
data — every other leaf leaves the record untouched. (This is the default mode;
AUTOMATICALLY_TRANSLATE replaces it — see Translation backfill.)
flowchart TD
RUN(["Run: for each configured model"]) --> PF{"Any localized SEO field has<br/>'required SEO fields' validation?"}
PF -->|"yes, and AUTOMATICALLY_TRANSLATE off"| ABORT[["🚫 ABORT: warn and exit<br/>(no records read or written)"]]
PF -->|"yes, but AUTOMATICALLY_TRANSLATE on"| REC(["For each record updated<br/>on/after the cutover date"])
PF -->|"no"| REC
REC --> CLS{"Classify SOURCE-locale SEO<br/>(title + description) with eld"}
CLS -->|"reliably the SOURCE language"| OK_SRC["✅ OK — already source locale"]
CLS -->|"reliably a FOREIGN language"| INV_C
CLS -->|"best-guess FOREIGN (low confidence)"| INV_S
CLS -->|"best-guess SOURCE (low confidence)"| OK_GUESS["✅ OK — left untouched<br/>(no positive corruption signal)"]
CLS -->|"empty"| EMPTY{"Any other locale<br/>populated?"}
CLS -->|"undetectable language"| UNK{"Any other locale<br/>populated?"}
EMPTY -->|"no"| OK_EMPTY["✅ OK — no SEO content"]
EMPTY -->|"yes"| INV_C
UNK -->|"no"| OK_UNK["✅ OK — undetectable & untranslated"]
UNK -->|"yes"| MANUAL["⚠️ Manual review<br/>(undetectable source, others populated)"]
subgraph INVESTIGATE ["investigate() — walk version history, newest first"]
direction TB
INV_C[["enter: tier = CONFIRMED"]] --> LISTV
INV_S[["enter: tier = SUSPECTED"]] --> LISTV
LISTV{"List item versions"}
LISTV -->|"read error"| ERR["🚫 Error reading history"]
LISTV -->|"no versions"| NOVER["⚠️ No version history"]
LISTV -->|"ok"| PRIM{"PRIMARY found?<br/>newest version whose SOURCE SEO<br/>is reliably the source language"}
PRIM -->|"yes"| TIER{"tier (from entry)"}
TIER -->|"CONFIRMED"| REV_OK["✅ Revert (primary) — WRITE"]
TIER -->|"SUSPECTED"| REV_WARN["⚠️ Revert (primary, low-confidence) — WRITE"]
PRIM -->|"no"| FB{"FALLBACK found?<br/>newest version: SOURCE populated,<br/>not foreign, ALL other locales empty"}
FB -->|"yes"| REV_FB["⚠️ Revert (fallback heuristic) — WRITE"]
FB -->|"no"| MOVEQ{"SOURCE content reliably foreign,<br/>its own locale exists and is blank?"}
MOVEQ -->|"yes"| MOVE["⚠️ Move title/description<br/>SOURCE → detected locale — WRITE"]
MOVEQ -->|"no"| NOGOOD["⚠️ No last-known-good version"]
end
Legend: ✅ = OK / no action ·
With AUTOMATICALLY_TRANSLATE = true and a DEEPL_API_TOKEN set, the goal
shifts from "fix the source locale" to "good enough" SEO in every locale,
still preferring real data over machine output. Per SEO field:
- Detect each locale's actual language (
eld, with a confidence score). - Keep content already correctly placed; move/revert misplaced-but- correct-language originals into their home locale with no re-translation — including reverting a known-good source from version history.
- Translate only the gaps: any locale still empty or detectably wrong is filled by DeepL-translating the highest-confidence content into it. Content already in the right language is never re-translated, and the source locale is never left holding a foreign value.
- If nothing can be reliably detected anywhere, the record is left for manual review.
Records resolved purely by keep/move/revert are ✅; any record that needed a
machine translation is title/description are translated
(image/twitter_card/no_index are preserved per locale), and DeepL output
is truncated to the field's own length validators. DeepL is called even under
DRY_RUN, so the log previews the exact translations it would write — this
uses DeepL quota.
All configuration lives in a gitignored .env file — copy .env.example to
.env and fill it in (loaded via dotenv).
Nothing project-specific is hardcoded in the source.
| Variable | Meaning |
|---|---|
DATOCMS_API_TOKEN |
CMA token: read records/versions + write records. |
DEEPL_API_TOKEN |
DeepL API key. Only needed when AUTOMATICALLY_TRANSLATE=true. Free keys end in :fx (free/pro endpoint auto-detected). |
ENVIRONMENT |
Environment to operate on (e.g. main or a sandbox). |
SOURCE_LOCALE |
Source locale to validate/restore, exactly as in the project (e.g. en). |
MODEL_IDS |
Comma-separated model (item type) IDs to check. |
UPDATED_SINCE |
Only scan records updated on/after this ISO timestamp (your cutover date). Blank = all records. |
DRY_RUN |
true (default) = report only, write nothing. false = apply changes. |
AUTOMATICALLY_TRANSLATE |
false (default) = revert/move only. true = also DeepL-backfill missing/incorrect locales (see Translation backfill). |
Optional knobs (sensible defaults): MAX_RECORDS_PER_MODEL,
MAX_VERSIONS_PER_RECORD, ADMIN_DOMAIN_OVERRIDE, and the rate/concurrency
settings under Parallelism & rate limits.
eld ships preloaded with a 60-language database; detection is automatically
restricted to your environment's locales.
cd ai-translations-seo-cleanup
npm install
cp .env.example .env # then fill in the values (DATOCMS_API_TOKEN, ENVIRONMENT, MODEL_IDS, ...)
# 1) Review with a dry run first (DRY_RUN=true in .env)
node cleanup.mjs
# 2) Inspect the generated seo-cleanup-log-<timestamp>.txt and seo-cleanup-results-<timestamp>.csv
# 3) Set DRY_RUN=false in .env and run again to apply
node cleanup.mjsRecords are processed concurrently, and each service is throttled by a
token-bucket limiter (the limiter
package) to stay within its rate limit:
| Setting | Default | Env override |
|---|---|---|
| DatoCMS requests/sec | 20 | DATOCMS_MAX_RPS |
| DeepL requests/sec | 20 | DEEPL_MAX_RPS |
| Records in parallel | 10 | MAX_CONCURRENCY |
DatoCMS 429s are additionally auto-retried by the CMA client; DeepL 429/529
responses are retried with exponential backoff (honoring Retry-After). See the
DatoCMS technical limits.
With concurrency, each record's log lines stay grouped, but records may appear
out of order (every line carries the record id).
A no-network smoke test exercises the detection, version-selection and
revert-payload logic against the real eld library:
node test-logic.mjs # or: npm testProgress is printed to the console and written to seo-cleanup-log-<timestamp>.txt.
Each record/field line is one of:
- ✅ OK — already in the source locale, no SEO content, or confidently reverted.
⚠️ Double-check — reverted via the weaker fallback heuristic, or corruption detected but no trustworthy version found (left untouched). Review these.- 🚫 Error — an API/processing error for that record (nothing written).
Each file entry includes the record's editing URL so you can jump straight to it.
Alongside the log, a seo-cleanup-results-<timestamp>.csv is written — one
row per changed record/field with the before/after title and description for
every locale, so the client can diff exactly what changed. Both files are
appended per record, so a run that fails or is aborted still leaves the work done
so far on disk.
Always run a dry run first and skim the log before setting
DRY_RUN=falsein.env. The tool is conservative — when in doubt it flags rather than writes — but you are editing production content. (WithAUTOMATICALLY_TRANSLATE=true, a dry run still calls DeepL, so the CSV and log preview the real translations.)