A hybrid AI operations platform for safe, end-to-end e-commerce post-purchase resolution
Natural-language understanding, deterministic business authority, selective agent orchestration, human approval, persistent reconciliation, and reliability evidence in one product loop.
ResolveOps turns a customer's natural-language request into a verified business outcome without giving an LLM unrestricted authority over consequential actions.
The product supports four post-purchase journeys—cancellation, address change, return, and refund—and makes the boundary between AI judgment and business control explicit:
The LLM interprets intent. The product verifies identity and policy. The router selects execution. Controlled tools change business state. Humans authorize exceptions.
Scope note: ResolveOps is a portfolio-grade prototype running on synthetic customers, orders, policies, and evaluation data. Reported results are controlled offline benchmark evidence, not production KPIs.
| Dimension | What ResolveOps provides |
|---|---|
| Customer problem | Resolve post-purchase requests without exposing internal workflow complexity |
| Supported intents | Cancellation, shipping-address change, return, refund |
| Customer experience | Natural-language request, guided entity confirmation, status and outcome reconciliation |
| Decision architecture | LLM understanding + product-owned authority + Workflow/Agent hybrid routing |
| Governance | Persistent Human Review, approve/reject controls, controlled post-approval execution |
| Reliability | Idempotent writes, state-machine checks, ownership re-verification, safe failure states |
| Operations | Reliability, risk, efficiency, trajectory, governance, and evidence-provenance dashboards |
| Evidence | Frozen linguistic, Agent, and orchestration evaluations with artifact integrity checks |
A request such as:
“I don't want to keep the headphones I bought recently. Can you help me send them back?”
contains several different problems that should not be collapsed into one prompt:
- Meaning: Is this a return, refund, cancellation, or an ambiguous combination?
- Identity: Which authenticated order and item does “the headphones” refer to?
- Authority: Is the item eligible under its current state, market policy, and risk level?
- Orchestration: Is the path already deterministic, or does it require dynamic planning?
- Execution: Did an authorized tool call actually commit the intended state change?
- Lifecycle: Is the customer's business outcome complete, pending, rejected, or awaiting review?
ResolveOps treats each question as a separate product layer so semantic confidence never becomes execution permission by accident.
| User | Job to be done | Product surface |
|---|---|---|
| Customer | Explain a problem naturally, identify the correct purchase, and understand what happened | Resolution Assistant |
| Operations reviewer | Inspect a consequential request, approve or reject it, and execute an approved action safely | Human Review |
| AI PM / Reliability lead | Compare orchestration choices, find failure patterns, and decide where Agent inference adds value | Reliability & Operations |
| Engineer / Evaluator | Reproduce tests, inspect traces, validate frozen evidence, and diagnose regressions | Test and evaluation harness |
The customer starts with an authenticated account and a natural-language request—no intent dropdown or required order ID. The assistant interprets the goal, finds candidate business objects, asks for confirmation when identity is not explicit, gathers any missing execution context, and exposes the final route and business outcome.
Human Review is a continuation of the business workflow, not an escalation dead end. Reviewers see the requested action, verified target, risk, policy rationale, and current approval state. They can approve or reject; approved work is then executed through the same controlled tools and reconciled back to the customer.
The internal control plane connects product decisions to evidence. It combines reliability KPIs, performance by intent and risk, tool and trajectory inspection, Human Review queue state, Workflow/Agent/Hybrid comparison, and SHA-256 provenance verification for frozen evaluation artifacts.
flowchart LR
A[Authenticated customer] --> B[Natural-language request]
B --> C[Understand intent and request shape]
C --> D[Discover candidate order/item]
D --> E{Identity explicit?}
E -->|No| F[Customer confirmation]
E -->|Yes| G[Ownership re-verification]
F --> G
G --> H[Collect and validate missing context]
H --> I[Product authority: state, policy, risk, permissions]
I --> J{Hybrid router}
J -->|Deterministic| K[Workflow]
J -->|Compositional / unresolved| L[Agent]
J -->|Approval required| M[Human Review]
J -->|Invalid / ineligible| N[Safe block]
K --> O[Controlled business tools]
L --> O
M --> O
O --> P[(PostgreSQL business state)]
P --> Q[Customer outcome reconciliation]
O --> R[Telemetry and evidence]
R --> S[Reliability & Operations]
| Stage | Product responsibility | Authority boundary | Observable output |
|---|---|---|---|
| Understand | Extract intent candidates, request shape, explicit IDs, product/time hints, and missing context | Model output is semantic evidence only | Structured understanding result |
| Identify | Search only the authenticated customer's orders/returns and rank candidates | Ranking does not establish identity | Candidate set with match evidence |
| Confirm | Bind the request to one order item and, for refunds, one return record | Candidate membership and ownership are re-verified | Confirmed business target |
| Validate | Parse a writable address, resolve return context, refresh order state, and evaluate policy | Raw model text cannot become a write payload | Complete execution context or a request for input |
| Route | Choose Workflow for resolved single-goal paths and Agent for compositional/unresolved orchestration | Routing selects strategy; it does not expand permissions | Workflow, Agent, Human Review, or safe block |
| Execute | Call only intent-scoped tools using deterministic rules and state transitions | Allowed actions are fixed in the runtime contract | Typed execution trace and persisted state |
| Reconcile | Read approval and business state after execution or review | Database state outranks stale UI state | Truthful customer-facing lifecycle status |
| Intent | Resolution path | Key safeguards | Typical terminal state |
|---|---|---|---|
| Cancellation | Verify order/item → load cancellation policy → validate order state → cancel | Ownership check, allowed-state policy, state-machine transition, idempotent repeat handling | Cancelled or safely rejected |
| Address change | Verify target → collect address → validate Singapore address structure → check order/shipment state → update | Model-extracted text is never written directly; concurrent state is checked at commit | Updated, needs input, or rejected |
| Return | Verify delivered item → load category/market policy → evaluate window/restriction → create return or request approval | Delivery date, return window, restricted-item policy, approval match, idempotency key | Return created, review pending, or rejected |
| Refund | Verify order/item → discover and confirm authenticated return → require completed return → evaluate amount threshold → initiate or review | Return ownership/state, amount bounds, exact approval payload match, idempotency protection | Refund pending, review pending, or rejected |
The interface exposes progression rather than pretending every request is immediately executable:
request_received
├─→ clarification_required
├─→ confirmation_required
├─→ additional_context_required
├─→ agent_planning_required
└─→ ready_for_execution
├─→ workflow
└─→ agent
Multi-intent or underspecified requests are clarified; ambiguous entities are confirmed; address and refund requests cannot execute until their additional authority inputs are complete.
The request-understanding layer handles natural phrasing, implicit goals, colloquial language, explicit identifiers, temporal hints, conditional requests, and multi-intent detection. It improves interface coverage but cannot authorize a business action.
ResolveOps keeps four concepts separate:
Intent understanding
≠ Target identity
≠ Policy eligibility
≠ Execution authorization
The authoritative context is rebuilt immediately before execution from PostgreSQL. Each intent receives a fixed capability set, such as get_order, search_policy, and only the relevant write tool. A router may choose an Agent, but it cannot grant the Agent additional capabilities.
| Condition | Preferred route | Rationale |
|---|---|---|
| Explicit or resolved single goal with complete context | Workflow | The next steps are known; deterministic execution is cheaper and easier to verify |
| Conditional, ambiguous, compositional, or state-dependent goal | Agent | Dynamic planning adds coverage where the next step is not fully predetermined |
| Policy requires authorization | Human Review | Consequential exceptions need persistent human permission |
| Ownership, state, policy, or input validation fails | Safe block / clarification | Uncertainty is surfaced instead of converted into action |
The core optimization target is:
Coverage × Reliability × Cost
The goal is not to replace workflows with Agents; it is to spend Agent inference only where it creates measurable value.
Business tools own policy checks, state transitions, database writes, and duplicate-action protection. Consequential return/refund writes use idempotency keys. Conditional updates guard against stale state. Tool results preserve distinctions among success, rejection, human review, not found, and infrastructure error.
Policy or risk gate
↓
PENDING approval
↓
APPROVED or REJECTED
↓
Controlled execution (approved only)
↓
Persistent business state
↓
Customer reconciliation
Three states remain deliberately separate:
Approval ≠ execution ≠ business outcome.
An approval authorizes one exact action and payload; it does not itself prove that execution succeeded. If execution fails after approval, the approval remains approved and the action can be retried safely. Idempotency prevents a retry from creating duplicate returns or refunds.
Two operating rules follow:
Persistent business state > UI session state.
AI task completion ≠ business lifecycle completion.
For example, “return request created” is not presented as “return completed,” and an approved refund is not presented as initiated until the refund record exists.
The operations surface is designed around diagnosis and product decisions, not vanity metrics.
| Area | What is measured or inspected |
|---|---|
| Reliability | Verified resolution, outcome correctness, evidence sufficiency, trajectory safety, truthful communication |
| Governance | Pending/approved/rejected reviews, executed-after-approval state, high-risk queue, oldest pending request |
| Segmentation | Performance by intent, risk level, and task stability |
| Efficiency | Tool calls, model calls, tokens, latency, repeated actions |
| Trajectory | Action distribution, write-before-read checks, per-trial tool sequence and variance |
| Architecture | Frozen Workflow-only vs Agent-only vs Hybrid comparison and route mix |
| Provenance | Manifest-backed artifact identity and SHA-256 verification before evidence is displayed |
If a frozen artifact fails integrity verification, the dashboard fails closed instead of presenting the comparison as trusted evidence.
ResolveOps uses an evaluation loop that checks both what happened and how it happened:
Define contract → freeze system/data → run evaluation → inspect outcome + trajectory
→ localize failure → change product/system → regression test → freeze evidence
On a frozen 25-case synthetic linguistic stress test:
| Understanding strategy | Strict semantic accuracy |
|---|---|
| Deterministic interpreter | 6 / 25 — 24% |
| LLM request understanding | 22 / 25 — 88% |
This is a semantic-layer result, not an end-to-end Agent uplift. It supports using an LLM at the interface while retaining deterministic business control.
The formal benchmark contains 16 tasks × 3 repetitions = 48 trials.
| Metric | Result |
|---|---|
| Strict task pass | 48 / 48 — 100% |
| Outcome correct | 100% |
| Trajectory safe | 100% |
| Evidence sufficient | 100% |
| Escalation appropriate | 100% |
| Communication truthful | 100% |
| Infrastructure failures | 0 |
| Average tool calls | 2.29 |
| Average model calls | 3.29 |
| Average tokens | 3,164 |
| Average latency | 7.62 s |
The final architecture decision was tested on 24 synthetic cases with unseen linguistic and orchestration conditions.
| Architecture | Strict pass | Actionable resolution | Unsafe cases |
|---|---|---|---|
| Workflow-only | 9 / 24 — 37.5% | 5 / 19 — 26.3% | 1 |
| Agent-only | 22 / 24 — 91.7% | 17 / 19 — 89.5% | 0 |
| Hybrid | 24 / 24 — 100% | 19 / 19 — 100% | 0 |
Hybrid route usage: 17 Workflow, 2 Agent, 5 no-execution.
| Metric | Agent-only | Hybrid | Change |
|---|---|---|---|
| Strict resolution | 91.7% | 100% | +8.3 pp |
| Average model calls | 4.08 | 1.29 | −68.4% |
| Average tokens | 6,039 | 3,413 | −43.5% |
| Average latency | 26.06 s | 21.69 s | −16.8% |
| Unsafe cases | 0 | 0 | No regression |
These results support the final product decision:
Preserve deterministic Workflow execution after semantic and business uncertainty are resolved; use Agent orchestration only when dynamic planning adds coverage.
All reported results are offline synthetic evidence. The repository tracks selected frozen artifacts under experiments/results/ and supporting audits under docs/, including:
- Evaluation strategy
- Evaluation and product decision
- Benchmark fairness and leakage audit
- v3A linguistic evidence
- v3B end-to-end evidence
- Final architecture and evidence
| Observed failure | Product risk | Design response |
|---|---|---|
| Early UI required intent and order IDs | Internal system structure was pushed onto the customer | Added natural-language understanding, authenticated discovery, candidate ranking, and explicit confirmation |
| Ranked candidate could be mistaken for confirmed identity | A plausible match could receive an unauthorized write | Made ranking descriptive only; confirmation and ownership verification establish identity |
| Approval existed without completed execution | Reviewers could believe the case was handled when business state had not changed | Rebuilt the flow as authorize → execute → reconcile |
| Frontend state became stale after Human Review | Customer messaging could diverge from the database | Reconcile against persistent approval and business state on return to the customer view |
| Agent execution was used on already-known paths | Higher latency and token cost without added coverage | Added the Hybrid router and preserved deterministic Workflow paths |
These failures remain represented in the regression suite rather than being hidden from the project history.
ResolveOps/
├── dashboard/ # Customer, reviewer, and reliability product surfaces
├── src/
│ ├── agents/ # Request understanding, Workflow baseline, Agent runtime
│ ├── environment/ # Database, rules, state machine, secrets bridge
│ ├── evaluation/ # Datasets, graders, runners, traces, usage metrics
│ ├── product/ # Entity resolution, conversation state, hybrid execution
│ ├── telemetry/
│ └── tools/ # Controlled read/write business capabilities
├── database/
│ ├── schema/
│ ├── migrations/
│ ├── seed/
│ └── dashboard_views_v1.sql
├── experiments/ # Agent and orchestration experiment runners + frozen results
├── docs/ # Architecture, decision logs, evaluation and audit evidence
├── scripts/
│ ├── bootstrap_local.py
│ └── bootstrap_remote.py
├── tests/ # Unit and integration regression coverage
├── docker-compose.yml
├── requirements.txt
└── .env.example
- Python 3.12
- Docker with Docker Compose
- An OpenAI-compatible model endpoint and API key
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txtcp .env.example .envSet local PostgreSQL credentials and the following model-provider values in .env:
LLM_API_KEY=your_api_key
LLM_BASE_URL=your_openai_compatible_base_url
LLM_MODEL=your_model_namepython scripts/bootstrap_local.py --check
python scripts/bootstrap_local.pyThe bootstrap sequence starts PostgreSQL, applies schema and migrations, loads synthetic business data and frozen evaluation telemetry, creates dashboard views, and runs a data-layer smoke test.
python scripts/bootstrap_local.py --reset-volumedeletes the local ResolveOps PostgreSQL volume before rebuilding it. Use it only when you intentionally want to discard the current demo state.
python -m streamlit run dashboard/app.pyOpen the local URL printed by Streamlit.
Hosted deployment uses a managed PostgreSQL provider because Streamlit Community Cloud does not run this repository's Docker Compose service.
Create a PostgreSQL database with a provider such as Neon or Supabase, export its connection values, then run:
export POSTGRES_HOST=your-project.region.aws.neon.tech
export POSTGRES_PORT=5432
export POSTGRES_DB=resolveops
export POSTGRES_USER=your_user
export POSTGRES_PASSWORD=your_password
export POSTGRES_SSLMODE=require
python scripts/bootstrap_remote.py --check
python scripts/bootstrap_remote.py--check validates files and connectivity without writing. The full command applies the idempotent initialization sequence and smoke test.
In Streamlit Community Cloud, create an app with:
Repository: lessismore4825-jl/ResolveOps
Branch: main
Main file: dashboard/app.py
Add the real values from .streamlit/secrets.toml.example to Streamlit Secrets, including POSTGRES_SSLMODE = "require", then deploy. src/environment/secrets_bridge.py maps hosted secrets to the same environment-variable contract used locally.
The hosted demo writes shared synthetic state and calls the configured LLM provider. Any visitor may mutate the demo data and consume API quota. Re-run
scripts/bootstrap_remote.pyto restore the frozen demo state.
python -m unittest discover -s tests -p 'test_*.py' -qThe suite covers business rules and state transitions, entity authority, Workflow and Agent behavior, Hybrid routing, Human Review, idempotent execution, persistent reconciliation, evaluation contracts, dashboard behavior, evidence integrity, and repository portability.
| Layer | Technology |
|---|---|
| Product and evaluation | Python 3.12 |
| Customer and operations UI | Streamlit |
| Business state | PostgreSQL 16 + psycopg 3 |
| Visualization | Plotly |
| Model integration | OpenAI-compatible API |
| Formal benchmark model | Qwen/Qwen3.5-35B-A3B |
| Local infrastructure | Docker Compose |
| Testing | Python unittest |
| Evidence integrity | SHA-256 manifest verification |
ResolveOps demonstrates AI product architecture, decision authority, controlled execution, human governance, evaluation, and reliability operations. It is not a production customer-service platform.
Current boundaries include:
- synthetic customers, orders, returns, policies, and evaluation datasets;
- offline benchmark results rather than production traffic or causal business impact;
- no production authentication, RBAC, SLO/SLA alerting, incident integration, or drift pipeline;
- no live payment, logistics, CRM, or customer-support integrations;
- no production CSAT, deflection, resolution-time, or cost-to-serve measurement;
- a deliberately bounded set of post-purchase intents and Singapore address validation.
A production rollout would require security and privacy review, real policy integration and versioning, authorization controls, audit retention, provider resilience, observability, calibrated human operations, and staged online evaluation.
ResolveOps expands the customer interface with AI while preserving deterministic business control.
The most important design decision is not the model. It is the separation of understanding, identity, eligibility, authorization, execution, and business outcome—with evidence available at every boundary.




