Skip to content
neon-x-hubPublic

About

A cost-based query optimizer and autonomous execution gateway for directed graph APIs.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

80 Commits

Folders and files

Repository files navigation

SpanQL

SpanQL Banner

SpanQL is a high-performance, isomorphic query compilation and execution framework for JavaScript and TypeScript. It enables a browser client to express a hierarchical data request as a compact URL, and a server gateway to decode, plan, concurrently resolve, and stitch that request into a nested JSON payload.

The framework is built around four concerns: a wire protocol that compresses queries into base64 bitmasks, a cost-based optimizer that generates parallel execution plans, a concurrent resolver engine with retry and backoff, and a stitching layer that merges flat resolver payloads back into nested JSON.


Packages

The monorepo ships three installable packages:

Package Environment Purpose
@spanql/ql-lib Isomorphic Wire protocol, GNS registry, DSL parser, schema compiler
@spanql/core Server (Node.js) Query planner, execution engine, stitcher, type validator
@spanql/client Browser / Node.js Query compiler and HTTP client

How It Works

A SpanQL request travels through five stages:

[Textual DSL Query]
        |
        v  @spanql/client
[SpanQLBrowserClient.encodeQueryParams()]
  - Parses DSL text into an AST
  - Resolves every attribute and modifier to a GNS slot index
  - Encodes selected slots into a base64 bitmask (q)
  - Encodes filter predicates and modifiers into a parameter map (p)
  - Identifies root entity as a context ID (r)
        |
        v  HTTP GET ?r=1&q=FGi&p=[7:"Dune"]
[GatewayQueryDecoder.decodeRequest()]   @spanql/core
  - Inflates the base64 bitmask back into an active slot list
  - Builds an AST QueryTree, classifying selections and modifiers
  - Translates the tree into executable QueryNode capsules
        |
        v
[QueryPlanner.generatePlan()]
  - Evaluates prune edges: marks nodes whose filters are satisfied
    by the active parameter set, short-circuiting their execution
  - Calculates expected row cardinalities per node using the
    Selectivity Factor Map (statistical models: uniform, histogram,
    parametric normal, discrete, uniform-branching-factor)
  - Orients DAG edges between active nodes based on cardinality
  - Returns an ExecutionPlan: entrypoints + dependency list
        |
        v
[ExecutionEngine.execute()]
  - Dispatches zero-dependency entrypoints immediately
  - As each resolver settles, decrements dependents' counts
    and triggers newly unblocked resolvers
  - Retries failed resolvers up to the configured count
  - Validates resolved records against the schema in strong mode
  - Captures per-node telemetry: queue time, execute time, count
        |
        v
[StitchingEngine.stitch()]
  - Post-order traversal: stitches deepest children first
  - Swaps foreign key references in parent records with resolved
    child objects (O(1) lookup by private_key)
  - Evaluates CEL filters, _offset, and _limit on pruned nodes
    whose records are already nested inside parent objects
  - Recursively strips attributes not requested in the query
        |
        v
[ExecutionReport] { data, plan, snapshot }

Installation

npm install @spanql/ql-lib @spanql/core
npm install @spanql/client

Schema Definition

Schemas are declared in YAML. Every entity must have exactly one field typed ID, which serves as its private key for stitching joins. Relationships are expressed by setting a field's type to another entity name. Wrapping the entity name in array brackets ([entity]) declares a has-many relationship. Appending ? to any type marks it as nullable.

# schema.spql.yml
author:
  id: ID
  name: string
  best_selling: book       # has-one relationship

book:
  id: ID
  title: string
  sales: number
  author: author           # back-reference (bidirectional)

Reserved field names that cannot be used in schemas: _limit, _offset, _filter.


Writing Resolvers

Resolvers are async functions that receive a ResolverContext and return a flat array of database records. The framework inspects the function source at initialization time to discover data dependency declarations.

const resolvers = {
  author: async (ctx) => {
    // Check whether the client requested the best_selling relation
    if (ctx.hasChild('best_selling')) {
      // Access parameters declared by the child node
      const { title } = ctx.child('best_selling').params();
      if (title.set) {
        // title.value holds the filter value; title.op.fn is the operator
        return await db.authors.where({ 'books.title': title.value });
      }
    }
    return await db.authors.findAll();
  },

  book: async (ctx) => {
    // Access parameters declared on this node directly
    const { sales } = ctx.params();
    return await db.books.findAll();
  }
};

Resolvers must return flat records where relational fields hold the foreign key ID of the related entity, not the nested object. The stitching engine performs the join.

// Correct: return flat records with foreign key IDs
{ id: 'a1', name: 'Frank Herbert', best_selling: 'b1' }

// Wrong: do not pre-nest objects in resolvers
{ id: 'a1', name: 'Frank Herbert', best_selling: { id: 'b1', ... } }

The exception is pruned nodes: if the planner determines a child node can be short-circuited because all its filter parameters are already satisfied by the parent resolver's query, the parent resolver is responsible for returning the nested object directly inside the parent record.


Server-Side Setup

import { SpanQLClient } from '@spanql/core';

const server = new SpanQLClient({
  schema: {
    file: './schema.spql.yml',         // path to YAML schema
  },
  resolvers: {
    retry: 2,                          // retry failed resolvers up to 2 times
    backoff: 100,                      // 100ms delay between retry attempts
    funcs: resolvers,                  // resolver function map
  },
  typing: {
    tolerance: 'strong',               // validate resolver output types
  },
  selectivity: {
    file: './stats.json',              // path to selectivity stats file
    refresh: 60000,                    // reload stats every 60 seconds
  },
});

await server.init();

// In your HTTP handler:
app.get('/spanql', async (req, res) => {
  const report = await server.query(req.query);
  res.json(report.data);
});

// On shutdown:
server.close();

SpanQLClientConfig options:

Option Type Default Description
schema.file string required Path to YAML schema file
typing.tolerance 'strong' | 'weak' 'strong' Type validation mode
typing.customs Record<string, ValidatorFn> — Custom type validators
resolvers.funcs Record<string, Function> required Resolver function map
resolvers.retry number 1 Max resolver retry attempts
resolvers.backoff number 50 Retry backoff in milliseconds
selectivity.file string — Path to selectivity stats JSON file
selectivity.refresh number — Stats file reload interval in milliseconds

Client-Side Setup

import { SpanQLBrowserClient } from '@spanql/client';

const client = new SpanQLBrowserClient({
  endpoint: 'https://api.my-app.com/spanql',
  schema: {
    author: { id: 'ID', name: 'string', best_selling: 'book' },
    book: { id: 'ID', title: 'string', sales: 'number', author: 'author' },
  },
  // Optional: override the GNS root context ID mapping.
  // By default, entities are assigned IDs by alphabetical sort order (1-based).
  // rootContextIds: { author: 1, book: 2 },
  //
  // Optional: inject a custom fetch implementation (e.g., for testing)
  // fetch: myCustomFetch,
});

Writing Queries

SpanQL queries use YAML-like indentation. Each line names a field. A colon followed by a value declares a filter predicate. A colon with no value opens a relationship block. Fields prefixed with _ are system modifiers.

author:
  name
  best_selling:
    title: "Dune"
    sales
    _filter: sales > 10000000
    _limit: 5
    _offset: 0

This query selects name on author, spans into best_selling, selects title (filtered to "Dune") and sales, then filters the result set by sales > 10000000, skips the first 0 records, and returns at most 5.

// Compile query to URL parameters
const url = client.buildQueryUrl(queryStr);
// https://api.my-app.com/spanql?r=1&q=FGi&p=[7:"Dune"]

// Or encode to { r, q, p } directly
const params = client.encodeQueryParams(queryStr);

// Or fire the HTTP request directly
const data = await client.query(queryStr);

Wire Protocol

Every SpanQL request carries three URL parameters:

Parameter Meaning Example
r Root context ID (integer identifying the root entity) 1
q Base64 topology bitmask: one bit per GNS slot index FGi
p Indexed parameter map: [slotIndex:value, ...] [7:"Dune"]

The GNS (Global Namespace) registry assigns a unique sequential slot index to every attribute, system modifier, and relationship in every entity up to a configurable traversal depth (default 4). Because the bitmask encodes presence by bit position rather than by name, only the root context ID r is needed to fully decode the topology on the server.


Query Modifiers

Three system modifiers are available on any entity or relationship block in a query:

Modifier Value Type Description
_filter CEL expression string Filter records using Google's Common Expression Language
_limit integer Maximum number of records to return
_offset integer Number of records to skip from the start

CEL expressions have access to all fields of each record as variables. String methods, arithmetic, comparisons, and boolean logic are fully supported.

books:
  title
  sales
  _filter: sales > 25000000 && title.startsWith("The")
  _limit: 10
  _offset: 20

Selection Pruning

Only fields explicitly requested in a query are present in the response. This includes relationship objects: if the query selects best_selling.title but not best_selling.id, the id field is stripped from all nested book objects in the output. Fields used internally for stitching joins (such as foreign key IDs) are preserved during resolution and removed in the final pruning pass.


Type Validation

In strong mode (the default), the execution engine validates every resolver's output records against the schema immediately after the resolver settles. Default validators cover string, number, boolean, and ID (which accepts both string and number). Custom validators can be registered for domain-specific types.

const server = new SpanQLClient({
  typing: {
    tolerance: 'strong',
    customs: {
      email:   (val) => typeof val === 'string' && val.includes('@'),
      percent: (val) => typeof val === 'number' && val >= 0 && val <= 100,
    },
  },
  // ...
});

Fields declared with a trailing ? (e.g., nickname: string?) are nullable; null and undefined values pass validation.


Selectivity Factor Map

The selectivity map is a JSON file that gives the query planner cardinality information per entity and per attribute. The planner uses this data to estimate expected row counts and orient execution DAG edges optimally.

{
  "author": {
    "total_rows": 10000,
    "attributes": {
      "name": { "type": "uniform", "params": { "sf": 0.001 } },
      "created_at": {
        "type": "equi-depth-histogram",
        "params": {
          "bins": [
            { "min": 1672531200, "max": 1680307200, "count": 3000 },
            { "min": 1680307200, "max": 1688083200, "count": 7000 }
          ]
        }
      }
    }
  }
}

Supported selectivity model types:

Model type string Use case
Uniform "uniform" Constant selectivity factor for evenly distributed values
Discrete "discrete" Map of categorical values to individual selectivity factors
Equi-depth histogram "equi-depth-histogram" Numeric range filters with non-uniform distributions
Parametric normal "parametric-normal" Numeric filters on normally distributed columns
Uniform branching factor "uniform-branching-factor" Average fan-out for has-many relationships

The file is optional. When omitted, all cardinalities default to 1 and DAG edges are oriented structurally rather than by cost.


Execution Report

SpanQLClient.query() returns an ExecutionReport:

interface ExecutionReport {
  data: any[];              // final nested JSON result array
  plan: ExecutionPlan;      // compiled task dependency graph
  snapshot: ExecutionSnapshot; // per-node runtime telemetry
}

interface ExecutionPlan {
  entrypoints: number[];    // node IDs with no dependencies (start first)
  tasks: Array<{
    node_id: number;
    type: string;           // entity name
    dependencies: number[]; // must settle before this task starts
    dependents: number[];   // triggered when this task settles
  }>;
}

interface ExecutionSnapshot {
  snapshot_id: string;
  total_execution_time_ms: number;
  node_metrics: Array<{
    node_id: number;
    status: 'SETTLED' | 'PRUNED';
    pruned_by_node_id: number | null;
    queued_at_ms: number;
    executed_at_ms: number;
    settled_at_ms: number;
    records_count: number;
    inbound_filter_keys_count: number;
  }>;
}

ResolverContext API

Every resolver function receives a ResolverContext instance as its sole argument.

Method / Property Signature Description
type string Entity name (e.g. "author")
node_id number Post-order traversal identifier
activePath string Dot-path from root (e.g. "author.best_selling")
status QueryNodeStatus Current lifecycle status
dependency_count number Remaining unsettled dependencies
pruned_by QueryNode | null Node that short-circuited this one, or null
params() Record<string, ASTValue> Active filter predicates for this node
results() any[] Currently resolved records array
setResults(r) void Overwrites the resolved records
hasChild(field) boolean Whether field was requested in the query
child(field) ResolverContext Context for a requested child relation (throws if absent)
childOrNull(field) ResolverContext | null Safe variant: returns null if child not requested
parent(path) ResolverContext Context for the parent node; validates active path boundary

ASTValue structure:

interface ASTValue {
  set: boolean;          // true if a filter value was provided by the client
  value: WireValue | WireValue[] | null;
  op: {
    fn: string;          // operator function name, e.g. "set", "gt", "rng", "sim"
    params: (string | number)[];
  };
}

Low-Level API: @spanql/ql-lib

The @spanql/ql-lib package is isomorphic and can be used independently of @spanql/core for direct wire format manipulation.

import {
  GNSRegistry,
  encodeQuery,
  decodeQuery,
  encodeTopology,
  encodeParameterMap,
  parseSpanQL,
  validateSchemaConfig,
} from '@spanql/ql-lib';

// Build a GNS registry from a schema map
const gns = new GNSRegistry(schemaConfigMap, { rootContextIds: { author: 1, book: 2 } });

// Look up a slot index by dot-path
const idx = gns.lookup('author.best_selling.title', 'author'); // => number

// Reverse lookup: slot index to path
const path = gns.revLookup(idx, 'author'); // => 'author.best_selling.title'

// Encode a query directly
const encoded = encodeQuery({ root: 'author', nodes: [{ path: 'author.name' }, { path: 'author.best_selling.title', value: 'Dune' }] }, gns);
// => { r: 1, q: 'FGi', p: '[7:"Dune"]' }

// Decode a query from wire parts
const tree = decodeQuery({ r: 1, q: 'FGi', p: '[7:"Dune"]' }, gns);
// => QueryTree

// Parse a DSL string into an AST
const dsl = parseSpanQL(`
author:
  name
  best_selling:
    title: "Dune"
`);

Documentation

About

A cost-based query optimizer and autonomous execution gateway for directed graph APIs.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages