Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

SEO locale cleanup tool

A standalone Node script that repairs DatoCMS records whose localized SEO field had its source locale overwritten by a target language.

Background

Versions <= 3.4.5 of the ai-translations plugin had a bug: translating a SEO field to more than one target locale at once could overwrite the source locale with one of the targets (e.g. an en -> es run wrote Spanish into the en slot). The plugin bug is fixed as of 3.4.6. This tool cleans up records that were already corrupted, bringing each one back to a last-known-good state of the source locale and blanking the other locales so they can be re-translated. Optionally, it can instead use DeepL to backfill every locale with "good enough" SEO in one pass (see Translation backfill).

It uses the DatoCMS CMA client to read/write records and version history, and eld (Efficient Language Detector) to decide which language a SEO value is actually in.

What it touches

  • Only the SEO field(s) of each record. No other field is ever read for writing or included in an update payload.
  • For each record of each configured model it:
    1. Reads the SEO field's source-locale title/description.
    2. Detects the language with eld, restricted to the environment's locales.
    3. If it is reliably the source language, leaves the record alone.
    4. Otherwise walks the record's version history to find the last known-good source-locale SEO and reverts only the source locale to it, while blanking the title/description of the other locales (their image / twitter_card / no_index are preserved).
  • Records are left as drafts — nothing is published.

Pre-flight safety check. Because step 4 blanks the title/description of non-source locales, the tool is incompatible with a SEO field that has the "required SEO fields" validation enabled for title or description: the API would reject every revert (VALIDATION_REQUIRED_SEO_FIELDS). The tool detects this from the schema and aborts before reading or writing anything, listing the offending field(s). To proceed, temporarily disable that validation (model settings → the SEO field → Validations), run the tool, then re-enable it.

Exception — AUTOMATICALLY_TRANSLATE. In translate mode the tool fills every locale rather than blanking it (keep / move / revert / DeepL-backfill), so the validator is no longer endangered. The check is therefore downgraded from an abort to a per-field warning and the run proceeds — no need to disable the validation. (One caveat: if a record's only known-good source genuinely lacks a required sub-field, that single write can still be rejected and is reported as a 🚫 for that record; the rest of the run is unaffected.)

How a "last-known-good" version is chosen

  1. Primary — the most recent version whose source SEO is reliably the source language.
  2. Fallback — if no version is reliably the source language, the most recent version whose source SEO is populated (and not reliably a foreign language) while every other locale is empty — i.e. the original, human-authored, pre-translation state. Reverts via the fallback are flagged Double-check.
  3. Move — if neither version exists but the source SEO is reliably a foreign language whose own locale is currently blank, the content is simply in the wrong place: its title/description are moved into that locale and cleared from the source. The reverse never happens — the source locale is never filled from a value detected in another locale. Flagged Double-check.
  4. If none of the above apply, the record is left untouched and flagged for manual review.

To avoid clobbering legitimate short titles, records whose source value is only low-confidence source language (too short for eld to be sure, but its best guess is the source language) are left untouched.

Decision flow

A schema pre-flight runs first (see the note above); if it passes, every record is taken through the tree below. Only the four WRITE leaves change data — every other leaf leaves the record untouched. (This is the default mode; AUTOMATICALLY_TRANSLATE replaces it — see Translation backfill.)

flowchart TD
  RUN(["Run: for each configured model"]) --> PF{"Any localized SEO field has<br/>'required SEO fields' validation?"}
  PF -->|"yes, and AUTOMATICALLY_TRANSLATE off"| ABORT[["🚫 ABORT: warn and exit<br/>(no records read or written)"]]
  PF -->|"yes, but AUTOMATICALLY_TRANSLATE on"| REC(["For each record updated<br/>on/after the cutover date"])
  PF -->|"no"| REC

  REC --> CLS{"Classify SOURCE-locale SEO<br/>(title + description) with eld"}

  CLS -->|"reliably the SOURCE language"| OK_SRC["✅ OK — already source locale"]
  CLS -->|"reliably a FOREIGN language"| INV_C
  CLS -->|"best-guess FOREIGN (low confidence)"| INV_S
  CLS -->|"best-guess SOURCE (low confidence)"| OK_GUESS["✅ OK — left untouched<br/>(no positive corruption signal)"]
  CLS -->|"empty"| EMPTY{"Any other locale<br/>populated?"}
  CLS -->|"undetectable language"| UNK{"Any other locale<br/>populated?"}

  EMPTY -->|"no"| OK_EMPTY["✅ OK — no SEO content"]
  EMPTY -->|"yes"| INV_C
  UNK -->|"no"| OK_UNK["✅ OK — undetectable & untranslated"]
  UNK -->|"yes"| MANUAL["⚠️ Manual review<br/>(undetectable source, others populated)"]

  subgraph INVESTIGATE ["investigate() — walk version history, newest first"]
    direction TB
    INV_C[["enter: tier = CONFIRMED"]] --> LISTV
    INV_S[["enter: tier = SUSPECTED"]] --> LISTV
    LISTV{"List item versions"}
    LISTV -->|"read error"| ERR["🚫 Error reading history"]
    LISTV -->|"no versions"| NOVER["⚠️ No version history"]
    LISTV -->|"ok"| PRIM{"PRIMARY found?<br/>newest version whose SOURCE SEO<br/>is reliably the source language"}
    PRIM -->|"yes"| TIER{"tier (from entry)"}
    TIER -->|"CONFIRMED"| REV_OK["✅ Revert (primary) — WRITE"]
    TIER -->|"SUSPECTED"| REV_WARN["⚠️ Revert (primary, low-confidence) — WRITE"]
    PRIM -->|"no"| FB{"FALLBACK found?<br/>newest version: SOURCE populated,<br/>not foreign, ALL other locales empty"}
    FB -->|"yes"| REV_FB["⚠️ Revert (fallback heuristic) — WRITE"]
    FB -->|"no"| MOVEQ{"SOURCE content reliably foreign,<br/>its own locale exists and is blank?"}
    MOVEQ -->|"yes"| MOVE["⚠️ Move title/description<br/>SOURCE → detected locale — WRITE"]
    MOVEQ -->|"no"| NOGOOD["⚠️ No last-known-good version"]
  end
Loading

Legend: ✅ = OK / no action · ⚠️ = written-but-verify, or manual review needed · 🚫 = error or aborted.

Translation backfill (optional)

With AUTOMATICALLY_TRANSLATE = true and a DEEPL_API_TOKEN set, the goal shifts from "fix the source locale" to "good enough" SEO in every locale, still preferring real data over machine output. Per SEO field:

  1. Detect each locale's actual language (eld, with a confidence score).
  2. Keep content already correctly placed; move/revert misplaced-but- correct-language originals into their home locale with no re-translation — including reverting a known-good source from version history.
  3. Translate only the gaps: any locale still empty or detectably wrong is filled by DeepL-translating the highest-confidence content into it. Content already in the right language is never re-translated, and the source locale is never left holding a foreign value.
  4. If nothing can be reliably detected anywhere, the record is left for manual review.

Records resolved purely by keep/move/revert are ✅; any record that needed a machine translation is ⚠️ (verify). Only title/description are translated (image/twitter_card/no_index are preserved per locale), and DeepL output is truncated to the field's own length validators. DeepL is called even under DRY_RUN, so the log previews the exact translations it would write — this uses DeepL quota.

Configuration

All configuration lives in a gitignored .env file — copy .env.example to .env and fill it in (loaded via dotenv). Nothing project-specific is hardcoded in the source.

Variable Meaning
DATOCMS_API_TOKEN CMA token: read records/versions + write records.
DEEPL_API_TOKEN DeepL API key. Only needed when AUTOMATICALLY_TRANSLATE=true. Free keys end in :fx (free/pro endpoint auto-detected).
ENVIRONMENT Environment to operate on (e.g. main or a sandbox).
SOURCE_LOCALE Source locale to validate/restore, exactly as in the project (e.g. en).
MODEL_IDS Comma-separated model (item type) IDs to check.
UPDATED_SINCE Only scan records updated on/after this ISO timestamp (your cutover date). Blank = all records.
DRY_RUN true (default) = report only, write nothing. false = apply changes.
AUTOMATICALLY_TRANSLATE false (default) = revert/move only. true = also DeepL-backfill missing/incorrect locales (see Translation backfill).

Optional knobs (sensible defaults): MAX_RECORDS_PER_MODEL, MAX_VERSIONS_PER_RECORD, ADMIN_DOMAIN_OVERRIDE, and the rate/concurrency settings under Parallelism & rate limits.

eld ships preloaded with a 60-language database; detection is automatically restricted to your environment's locales.

Run

cd ai-translations-seo-cleanup
npm install
cp .env.example .env        # then fill in the values (DATOCMS_API_TOKEN, ENVIRONMENT, MODEL_IDS, ...)
# 1) Review with a dry run first (DRY_RUN=true in .env)
node cleanup.mjs
# 2) Inspect the generated seo-cleanup-log-<timestamp>.txt and seo-cleanup-results-<timestamp>.csv
# 3) Set DRY_RUN=false in .env and run again to apply
node cleanup.mjs

Parallelism & rate limits

Records are processed concurrently, and each service is throttled by a token-bucket limiter (the limiter package) to stay within its rate limit:

Setting Default Env override
DatoCMS requests/sec 20 DATOCMS_MAX_RPS
DeepL requests/sec 20 DEEPL_MAX_RPS
Records in parallel 10 MAX_CONCURRENCY

DatoCMS 429s are additionally auto-retried by the CMA client; DeepL 429/529 responses are retried with exponential backoff (honoring Retry-After). See the DatoCMS technical limits. With concurrency, each record's log lines stay grouped, but records may appear out of order (every line carries the record id).

Verify the logic (optional)

A no-network smoke test exercises the detection, version-selection and revert-payload logic against the real eld library:

node test-logic.mjs   # or: npm test

Output

Progress is printed to the console and written to seo-cleanup-log-<timestamp>.txt. Each record/field line is one of:

  • ✅ OK — already in the source locale, no SEO content, or confidently reverted.
  • ⚠️ Double-check — reverted via the weaker fallback heuristic, or corruption detected but no trustworthy version found (left untouched). Review these.
  • 🚫 Error — an API/processing error for that record (nothing written).

Each file entry includes the record's editing URL so you can jump straight to it.

Alongside the log, a seo-cleanup-results-<timestamp>.csv is written — one row per changed record/field with the before/after title and description for every locale, so the client can diff exactly what changed. Both files are appended per record, so a run that fails or is aborted still leaves the work done so far on disk.

Always run a dry run first and skim the log before setting DRY_RUN=false in .env. The tool is conservative — when in doubt it flags rather than writes — but you are editing production content. (With AUTOMATICALLY_TRANSLATE=true, a dry run still calls DeepL, so the CSV and log preview the real translations.)

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages