Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
89 changes: 89 additions & 0 deletions docs/identifier-policy.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,89 @@
# Identifier policy

This page is the canonical description of how MCPg treats schema, table,
column, role, and other names. Tool descriptions and error messages should
stay consistent with it so agents can choose tools and recover from rejections.

MCPg follows a **match the sink** rule: naming limits depend on where the
value is used, not on a single global "safe name" style.

## SQL and `pg_dump` / `pg_restore` paths

Tools that address real PostgreSQL objects (query, export, import, dump,
replication, roles used in `SET ROLE`, and similar) accept any **addressable**
PostgreSQL identifier. Values are encoded with `quote_identifier` (or dump
pattern encoding for `pg_dump`), so injection is prevented by **escaping**,
not by forbidding hyphens.
Comment on lines +12 to +16

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue: The document claims SQL-oriented tools accept any addressable PostgreSQL identifier and tells agents to retry SQL tools with the same name, but NL→SQL deliberately rejects names requiring delimited quoting and the documented run_select example is not an available MCPg tool. An agent following this policy either gets another rejection or attempts a nonexistent tool.

Triggers: When an agent uses NL→SQL or follows the documented retry list after a plain-identifier rejection.

Suggested fix: List NL→SQL among the deliberate strict exceptions with its prompt-injection rationale, and replace run_select with an actual registered SQL tool name.


| Example | Notes |
|---------|--------|
| `adm-pgbench` | Hyphen — common real schema name (issue #329) |
| `My Schema` | Space |
| `Users` | Mixed case preserved only when quoted |
| `2fa_tokens` | Leading digit requires quoting in SQL |
| `app$cfg` | `$` is allowed in PostgreSQL unquoted identifiers; still fine when quoted |
| `données` | Non-ASCII letters |
| `order` / `user` | Reserved words — must be quoted as identifiers |
| `weird"name` | Embedded `"` is doubled to `""` inside quotes |

**Always rejected** (not valid addressable identifiers): empty string, embedded
NUL (`\x00`), names longer than **63 bytes** (PostgreSQL's `NAMEDATALEN - 1`;
the server would otherwise **silently truncate**).
Comment on lines +29 to +31

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

issue: The policy says names longer than 63 bytes are always rejected on SQL and pg_dump paths, but pg_dump pattern encoding accepts overlong names and passes them to pg_dump; callers therefore receive contradictory guidance about whether such names are valid for dump operations.

Triggers: When a caller uses an over-63-byte schema or table name with a dump tool.

Suggested fix: Either enforce the 63-byte limit in _encode_schema_pattern or document the dump-specific behavior separately from SQL identifier validation.

Suggested change
**Always rejected** (not valid addressable identifiers): empty string, embedded
NUL (`\x00`), names longer than **63 bytes** (PostgreSQL's `NAMEDATALEN - 1`;
the server would otherwise **silently truncate**).
**Always rejected for SQL identifier paths** (not valid addressable identifiers):
empty string, embedded NUL (`\x00`), names longer than **63 bytes**
(PostgreSQL's `NAMEDATALEN - 1`; the server would otherwise **silently
truncate**). `pg_dump` pattern paths may accept overlong names and pass them
through to `pg_dump`.


Callers pass the **bare** name (e.g. `adm-pgbench`). Do not add SQL quotes
yourself.

Beyond the hyphen, the important cases are spaces, case preservation, leading
digits, `$`, Unicode letters, SQL reserved words used as names, and embedded
double quotes — all legal in PostgreSQL delimited identifiers when encoded
correctly.

## Plain-identifier-only tools (deliberate)

Some tools feed the name into a **different** language or channel. Those keep
a stricter allowlist, typically `[A-Za-z_][A-Za-z0-9_]*`:

| Sink | Examples | Why |
|------|----------|-----|
| Generated source | Prisma, Drizzle, Diesel, Ecto, Ent, jOOQ, sqlc, SQLAlchemy export | Name becomes an identifier in TypeScript / Go / Rust / Elixir / Python |
| Graph labels | Apache AGE / graph projection labels | Extension label rules, not PG delimited identifiers |
| Shell / host command | `schedule_logical_backup` paths, some `COPY TO PROGRAM` args | Shell metacharacters are a real injection surface |
Comment on lines +48 to +50

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nitpick: The shell row describes schedule_logical_backup paths as plain-identifier-only inputs, but those inputs are path values with a different allowlist: destinations require absolute POSIX paths and pg_dump_path permits path separators, while the database argument accepts hyphens. Treating these as [A-Za-z_][A-Za-z0-9_]* names gives agents the wrong validation and recovery instructions.

Triggers: When an agent chooses or validates arguments for schedule_logical_backup.

Suggested fix: Describe shell/path arguments using their actual path-specific allowlists, and distinguish them from identifier validation for the database argument.

| Some SQL *literals* | `regconfig` in text-search helpers | Embedded as a string literal, not as a delimited identifier |
| Config keys | `MCPG_SECONDARY_DATABASE_URLS` names | MCPg's own ids (`[a-z0-9_]+`), not PG object names |

If such a tool rejects a name that SQL tools accept, the error should state the
**sink** and point at SQL/dump tools for the real object name.

### Suggested error shape

```text
PrismaExportError: invalid schema name 'adm-pgbench'; this tool only accepts
plain identifiers [A-Za-z_][A-Za-z0-9_]* because the name becomes a source
identifier in generated Prisma schema. Use a plain-named schema, or map the
PostgreSQL name in your own layer. SQL tools (export_table, dump_database, …)
accept delimited names via quoting.
```

### Suggested tool-description blurb

```text
Names must be plain PostgreSQL identifiers: [A-Za-z_][A-Za-z0-9_]*.
Hyphens, spaces, and other delimited-identifier characters are not supported
in this tool because the value is used as <SINK>, not only as a SQL identifier.
See docs/identifier-policy.md.
```

## What agents should do

1. Prefer tool descriptions and error text over assumptions from the README alone.
2. On a plain-identifier rejection, retry with SQL-oriented tools (`export_table`,
`dump_database`, `run_select`, …) using the same bare name.
3. Do not strip hyphens or rename production schemas solely to satisfy an ORM
exporter — map names in the generator layer if you need both.

## README note

The main README **Why MCPg** bullet should not claim that all identifier
interpolation still uses a global `[A-Za-z_][A-Za-z0-9_]*` allowlist. That was
true historically; since 0.8.2 / issue #329, SQL and `pg_dump` paths use
encoding. Link here from the README when that bullet is updated.
3 changes: 3 additions & 0 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,9 @@ them" — pick the section that matches what you're trying to do.

- [**Tools**](tools.md) — every MCP tool MCPg exposes, including
the capability gates that need to be on.
- [**Identifier policy**](identifier-policy.md) — which names SQL tools
accept (including hyphens and other delimited identifiers) versus
plain-identifier-only sinks (ORM exporters, AGE labels, shell paths).
- [**Architecture**](architecture.md) — how the pieces fit together
(server, drivers, replicas, cursors, audit, transports).
- [**Scaling guide**](scaling.md) — pool sizing, replica fan-out,
Expand Down
Loading