Skip to content

Repository files navigation

ResolveOps

A hybrid AI operations platform for safe, end-to-end e-commerce post-purchase resolution
Natural-language understanding, deterministic business authority, selective agent orchestration, human approval, persistent reconciliation, and reliability evidence in one product loop.

ResolveOps turns a customer's natural-language request into a verified business outcome without giving an LLM unrestricted authority over consequential actions.

The product supports four post-purchase journeys—cancellation, address change, return, and refund—and makes the boundary between AI judgment and business control explicit:

The LLM interprets intent. The product verifies identity and policy. The router selects execution. Controlled tools change business state. Humans authorize exceptions.

Scope note: ResolveOps is a portfolio-grade prototype running on synthetic customers, orders, policies, and evaluation data. Reported results are controlled offline benchmark evidence, not production KPIs.


Product at a Glance

Dimension What ResolveOps provides
Customer problem Resolve post-purchase requests without exposing internal workflow complexity
Supported intents Cancellation, shipping-address change, return, refund
Customer experience Natural-language request, guided entity confirmation, status and outcome reconciliation
Decision architecture LLM understanding + product-owned authority + Workflow/Agent hybrid routing
Governance Persistent Human Review, approve/reject controls, controlled post-approval execution
Reliability Idempotent writes, state-machine checks, ownership re-verification, safe failure states
Operations Reliability, risk, efficiency, trajectory, governance, and evidence-provenance dashboards
Evidence Frozen linguistic, Agent, and orchestration evaluations with artifact integrity checks

Why this product exists

A request such as:

“I don't want to keep the headphones I bought recently. Can you help me send them back?”

contains several different problems that should not be collapsed into one prompt:

  1. Meaning: Is this a return, refund, cancellation, or an ambiguous combination?
  2. Identity: Which authenticated order and item does “the headphones” refer to?
  3. Authority: Is the item eligible under its current state, market policy, and risk level?
  4. Orchestration: Is the path already deterministic, or does it require dynamic planning?
  5. Execution: Did an authorized tool call actually commit the intended state change?
  6. Lifecycle: Is the customer's business outcome complete, pending, rejected, or awaiting review?

ResolveOps treats each question as a separate product layer so semantic confidence never becomes execution permission by accident.


Users and Jobs to Be Done

User Job to be done Product surface
Customer Explain a problem naturally, identify the correct purchase, and understand what happened Resolution Assistant
Operations reviewer Inspect a consequential request, approve or reject it, and execute an approved action safely Human Review
AI PM / Reliability lead Compare orchestration choices, find failure patterns, and decide where Agent inference adds value Reliability & Operations
Engineer / Evaluator Reproduce tests, inspect traces, validate frozen evidence, and diagnose regressions Test and evaluation harness

Product Surfaces

1. Resolution Assistant

The customer starts with an authenticated account and a natural-language request—no intent dropdown or required order ID. The assistant interprets the goal, finds candidate business objects, asks for confirmation when identity is not explicit, gathers any missing execution context, and exposes the final route and business outcome.

Natural-language resolution
Move from a customer request to an observable, governed resolution state.

ResolveOps Resolution Assistant
Entity authority
Candidate ranking supports discovery; explicit confirmation establishes target identity.

ResolveOps entity confirmation

2. Human Review

Human Review is a continuation of the business workflow, not an escalation dead end. Reviewers see the requested action, verified target, risk, policy rationale, and current approval state. They can approve or reject; approved work is then executed through the same controlled tools and reconciled back to the customer.

3. Reliability & Operations

The internal control plane connects product decisions to evidence. It combines reliability KPIs, performance by intent and risk, tool and trajectory inspection, Human Review queue state, Workflow/Agent/Hybrid comparison, and SHA-256 provenance verification for frozen evaluation artifacts.

Governed execution
High-risk requests continue through persistent authorization and controlled execution.

ResolveOps Human Review
Reliability control plane
Outcome, safety, efficiency, governance, and evidence are visible in one operational surface.

ResolveOps Reliability and Operations

End-to-End Resolution Flow

flowchart LR
    A[Authenticated customer] --> B[Natural-language request]
    B --> C[Understand intent and request shape]
    C --> D[Discover candidate order/item]
    D --> E{Identity explicit?}
    E -->|No| F[Customer confirmation]
    E -->|Yes| G[Ownership re-verification]
    F --> G
    G --> H[Collect and validate missing context]
    H --> I[Product authority: state, policy, risk, permissions]
    I --> J{Hybrid router}
    J -->|Deterministic| K[Workflow]
    J -->|Compositional / unresolved| L[Agent]
    J -->|Approval required| M[Human Review]
    J -->|Invalid / ineligible| N[Safe block]
    K --> O[Controlled business tools]
    L --> O
    M --> O
    O --> P[(PostgreSQL business state)]
    P --> Q[Customer outcome reconciliation]
    O --> R[Telemetry and evidence]
    R --> S[Reliability & Operations]
Loading

Stage-by-stage contract

Stage Product responsibility Authority boundary Observable output
Understand Extract intent candidates, request shape, explicit IDs, product/time hints, and missing context Model output is semantic evidence only Structured understanding result
Identify Search only the authenticated customer's orders/returns and rank candidates Ranking does not establish identity Candidate set with match evidence
Confirm Bind the request to one order item and, for refunds, one return record Candidate membership and ownership are re-verified Confirmed business target
Validate Parse a writable address, resolve return context, refresh order state, and evaluate policy Raw model text cannot become a write payload Complete execution context or a request for input
Route Choose Workflow for resolved single-goal paths and Agent for compositional/unresolved orchestration Routing selects strategy; it does not expand permissions Workflow, Agent, Human Review, or safe block
Execute Call only intent-scoped tools using deterministic rules and state transitions Allowed actions are fixed in the runtime contract Typed execution trace and persisted state
Reconcile Read approval and business state after execution or review Database state outranks stale UI state Truthful customer-facing lifecycle status

Supported intent playbooks

Intent Resolution path Key safeguards Typical terminal state
Cancellation Verify order/item → load cancellation policy → validate order state → cancel Ownership check, allowed-state policy, state-machine transition, idempotent repeat handling Cancelled or safely rejected
Address change Verify target → collect address → validate Singapore address structure → check order/shipment state → update Model-extracted text is never written directly; concurrent state is checked at commit Updated, needs input, or rejected
Return Verify delivered item → load category/market policy → evaluate window/restriction → create return or request approval Delivery date, return window, restricted-item policy, approval match, idempotency key Return created, review pending, or rejected
Refund Verify order/item → discover and confirm authenticated return → require completed return → evaluate amount threshold → initiate or review Return ownership/state, amount bounds, exact approval payload match, idempotency protection Refund pending, review pending, or rejected

Conversational state model

The interface exposes progression rather than pretending every request is immediately executable:

request_received
    ├─→ clarification_required
    ├─→ confirmation_required
    ├─→ additional_context_required
    ├─→ agent_planning_required
    └─→ ready_for_execution
             ├─→ workflow
             └─→ agent

Multi-intent or underspecified requests are clarified; ambiguous entities are confirmed; address and refund requests cannot execute until their additional authority inputs are complete.


Decision Architecture

1. Semantic intelligence

The request-understanding layer handles natural phrasing, implicit goals, colloquial language, explicit identifiers, temporal hints, conditional requests, and multi-intent detection. It improves interface coverage but cannot authorize a business action.

2. Product authority

ResolveOps keeps four concepts separate:

Intent understanding
≠ Target identity
≠ Policy eligibility
≠ Execution authorization

The authoritative context is rebuilt immediately before execution from PostgreSQL. Each intent receives a fixed capability set, such as get_order, search_policy, and only the relevant write tool. A router may choose an Agent, but it cannot grant the Agent additional capabilities.

3. Hybrid orchestration

Condition Preferred route Rationale
Explicit or resolved single goal with complete context Workflow The next steps are known; deterministic execution is cheaper and easier to verify
Conditional, ambiguous, compositional, or state-dependent goal Agent Dynamic planning adds coverage where the next step is not fully predetermined
Policy requires authorization Human Review Consequential exceptions need persistent human permission
Ownership, state, policy, or input validation fails Safe block / clarification Uncertainty is surfaced instead of converted into action

The core optimization target is:

Coverage × Reliability × Cost

The goal is not to replace workflows with Agents; it is to spend Agent inference only where it creates measurable value.

4. Controlled execution

Business tools own policy checks, state transitions, database writes, and duplicate-action protection. Consequential return/refund writes use idempotency keys. Conditional updates guard against stale state. Tool results preserve distinctions among success, rejection, human review, not found, and infrastructure error.


Human-in-the-Loop Governance

Policy or risk gate
      ↓
PENDING approval
      ↓
APPROVED or REJECTED
      ↓
Controlled execution (approved only)
      ↓
Persistent business state
      ↓
Customer reconciliation

Three states remain deliberately separate:

Approval ≠ execution ≠ business outcome.

An approval authorizes one exact action and payload; it does not itself prove that execution succeeded. If execution fails after approval, the approval remains approved and the action can be retried safely. Idempotency prevents a retry from creating duplicate returns or refunds.

Two operating rules follow:

Persistent business state > UI session state.
AI task completion ≠ business lifecycle completion.

For example, “return request created” is not presented as “return completed,” and an approved refund is not presented as initiated until the refund record exists.


Reliability & Operations

The operations surface is designed around diagnosis and product decisions, not vanity metrics.

Area What is measured or inspected
Reliability Verified resolution, outcome correctness, evidence sufficiency, trajectory safety, truthful communication
Governance Pending/approved/rejected reviews, executed-after-approval state, high-risk queue, oldest pending request
Segmentation Performance by intent, risk level, and task stability
Efficiency Tool calls, model calls, tokens, latency, repeated actions
Trajectory Action distribution, write-before-read checks, per-trial tool sequence and variance
Architecture Frozen Workflow-only vs Agent-only vs Hybrid comparison and route mix
Provenance Manifest-backed artifact identity and SHA-256 verification before evidence is displayed

If a frozen artifact fails integrity verification, the dashboard fails closed instead of presenting the comparison as trusted evidence.


Evaluation and Product Decision

ResolveOps uses an evaluation loop that checks both what happened and how it happened:

Define contract → freeze system/data → run evaluation → inspect outcome + trajectory
→ localize failure → change product/system → regression test → freeze evidence

1. Linguistic generalization (v3A)

On a frozen 25-case synthetic linguistic stress test:

Understanding strategy Strict semantic accuracy
Deterministic interpreter 6 / 25 — 24%
LLM request understanding 22 / 25 — 88%

This is a semantic-layer result, not an end-to-end Agent uplift. It supports using an LLM at the interface while retaining deterministic business control.

2. Agent reliability (v1.2)

The formal benchmark contains 16 tasks × 3 repetitions = 48 trials.

Metric Result
Strict task pass 48 / 48 — 100%
Outcome correct 100%
Trajectory safe 100%
Evidence sufficient 100%
Escalation appropriate 100%
Communication truthful 100%
Infrastructure failures 0
Average tool calls 2.29
Average model calls 3.29
Average tokens 3,164
Average latency 7.62 s

3. Frozen orchestration holdout (v3B)

The final architecture decision was tested on 24 synthetic cases with unseen linguistic and orchestration conditions.

Architecture Strict pass Actionable resolution Unsafe cases
Workflow-only 9 / 24 — 37.5% 5 / 19 — 26.3% 1
Agent-only 22 / 24 — 91.7% 17 / 19 — 89.5% 0
Hybrid 24 / 24 — 100% 19 / 19 — 100% 0

Hybrid route usage: 17 Workflow, 2 Agent, 5 no-execution.

Metric Agent-only Hybrid Change
Strict resolution 91.7% 100% +8.3 pp
Average model calls 4.08 1.29 −68.4%
Average tokens 6,039 3,413 −43.5%
Average latency 26.06 s 21.69 s −16.8%
Unsafe cases 0 0 No regression

These results support the final product decision:

Preserve deterministic Workflow execution after semantic and business uncertainty are resolved; use Agent orchestration only when dynamic planning adds coverage.

ResolveOps Hybrid evaluation evidence

Evidence scope

All reported results are offline synthetic evidence. The repository tracks selected frozen artifacts under experiments/results/ and supporting audits under docs/, including:


Failure-Driven Product Iteration

Observed failure Product risk Design response
Early UI required intent and order IDs Internal system structure was pushed onto the customer Added natural-language understanding, authenticated discovery, candidate ranking, and explicit confirmation
Ranked candidate could be mistaken for confirmed identity A plausible match could receive an unauthorized write Made ranking descriptive only; confirmation and ownership verification establish identity
Approval existed without completed execution Reviewers could believe the case was handled when business state had not changed Rebuilt the flow as authorize → execute → reconcile
Frontend state became stale after Human Review Customer messaging could diverge from the database Reconcile against persistent approval and business state on return to the customer view
Agent execution was used on already-known paths Higher latency and token cost without added coverage Added the Hybrid router and preserved deterministic Workflow paths

These failures remain represented in the regression suite rather than being hidden from the project history.


Repository Structure

ResolveOps/
├── dashboard/                 # Customer, reviewer, and reliability product surfaces
├── src/
│   ├── agents/                # Request understanding, Workflow baseline, Agent runtime
│   ├── environment/           # Database, rules, state machine, secrets bridge
│   ├── evaluation/            # Datasets, graders, runners, traces, usage metrics
│   ├── product/               # Entity resolution, conversation state, hybrid execution
│   ├── telemetry/
│   └── tools/                 # Controlled read/write business capabilities
├── database/
│   ├── schema/
│   ├── migrations/
│   ├── seed/
│   └── dashboard_views_v1.sql
├── experiments/               # Agent and orchestration experiment runners + frozen results
├── docs/                      # Architecture, decision logs, evaluation and audit evidence
├── scripts/
│   ├── bootstrap_local.py
│   └── bootstrap_remote.py
├── tests/                     # Unit and integration regression coverage
├── docker-compose.yml
├── requirements.txt
└── .env.example

Run Locally

Prerequisites

  • Python 3.12
  • Docker with Docker Compose
  • An OpenAI-compatible model endpoint and API key

1. Create an environment and install dependencies

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

2. Configure the environment

cp .env.example .env

Set local PostgreSQL credentials and the following model-provider values in .env:

LLM_API_KEY=your_api_key
LLM_BASE_URL=your_openai_compatible_base_url
LLM_MODEL=your_model_name

3. Validate and bootstrap

python scripts/bootstrap_local.py --check
python scripts/bootstrap_local.py

The bootstrap sequence starts PostgreSQL, applies schema and migrations, loads synthetic business data and frozen evaluation telemetry, creates dashboard views, and runs a data-layer smoke test.

python scripts/bootstrap_local.py --reset-volume deletes the local ResolveOps PostgreSQL volume before rebuilding it. Use it only when you intentionally want to discard the current demo state.

4. Start the application

python -m streamlit run dashboard/app.py

Open the local URL printed by Streamlit.


Deploy to Streamlit Community Cloud

Hosted deployment uses a managed PostgreSQL provider because Streamlit Community Cloud does not run this repository's Docker Compose service.

1. Provision and initialize PostgreSQL

Create a PostgreSQL database with a provider such as Neon or Supabase, export its connection values, then run:

export POSTGRES_HOST=your-project.region.aws.neon.tech
export POSTGRES_PORT=5432
export POSTGRES_DB=resolveops
export POSTGRES_USER=your_user
export POSTGRES_PASSWORD=your_password
export POSTGRES_SSLMODE=require

python scripts/bootstrap_remote.py --check
python scripts/bootstrap_remote.py

--check validates files and connectivity without writing. The full command applies the idempotent initialization sequence and smoke test.

2. Deploy the app

In Streamlit Community Cloud, create an app with:

Repository: lessismore4825-jl/ResolveOps
Branch:     main
Main file:  dashboard/app.py

Add the real values from .streamlit/secrets.toml.example to Streamlit Secrets, including POSTGRES_SSLMODE = "require", then deploy. src/environment/secrets_bridge.py maps hosted secrets to the same environment-variable contract used locally.

The hosted demo writes shared synthetic state and calls the configured LLM provider. Any visitor may mutate the demo data and consume API quota. Re-run scripts/bootstrap_remote.py to restore the frozen demo state.


Test Suite

python -m unittest discover -s tests -p 'test_*.py' -q

The suite covers business rules and state transitions, entity authority, Workflow and Agent behavior, Hybrid routing, Human Review, idempotent execution, persistent reconciliation, evaluation contracts, dashboard behavior, evidence integrity, and repository portability.


Technology

Layer Technology
Product and evaluation Python 3.12
Customer and operations UI Streamlit
Business state PostgreSQL 16 + psycopg 3
Visualization Plotly
Model integration OpenAI-compatible API
Formal benchmark model Qwen/Qwen3.5-35B-A3B
Local infrastructure Docker Compose
Testing Python unittest
Evidence integrity SHA-256 manifest verification

Boundaries and Limitations

ResolveOps demonstrates AI product architecture, decision authority, controlled execution, human governance, evaluation, and reliability operations. It is not a production customer-service platform.

Current boundaries include:

  • synthetic customers, orders, returns, policies, and evaluation datasets;
  • offline benchmark results rather than production traffic or causal business impact;
  • no production authentication, RBAC, SLO/SLA alerting, incident integration, or drift pipeline;
  • no live payment, logistics, CRM, or customer-support integrations;
  • no production CSAT, deflection, resolution-time, or cost-to-serve measurement;
  • a deliberately bounded set of post-purchase intents and Singapore address validation.

A production rollout would require security and privacy review, real policy integration and versioning, authorization controls, audit retention, provider resilience, observability, calibrated human operations, and staged online evaluation.


Core Product Thesis

ResolveOps expands the customer interface with AI while preserving deterministic business control.

The most important design decision is not the model. It is the separation of understanding, identity, eligibility, authorization, execution, and business outcome—with evidence available at every boundary.

About

Hybrid AI Agent platform for governed e-commerce post-purchase resolution, with Product Authority, HITL, evaluation, and reliability operations.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages