Portable prompt assets meet portable eval assets — a Prompty ↔ EvalPort connection? #500
adhabnr-ux
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi Prompty maintainers,
I maintain EvalPort, an open spec (Apache 2.0) for portable LLM evaluation test suites, test cases, and result sets — JSON Schemas plus Python/TypeScript SDKs, aimed at letting eval data move between frameworks (DeepEval, Promptfoo, Inspect AI, LangSmith, Braintrust, etc.) without losing semantic fidelity.
I read through
spec/spec.md, theschema/TypeSpec model, and the README before posting, and the shape of what Prompty is solving struck me as a genuine sibling problem to what EvalPort solves — just one layer up the pipeline:.promptyfile (YAML frontmatter + markdown body) runs unchanged across Python, TypeScript, C#, Go, Rust, Java, and Swift runtimes, because the model is defined once in TypeSpec (schema/model/) and generated out to every language.Right now there's no link between the two — a
.promptyfile has no way to say "here's the eval suite this should be graded against," and an EvalPortResultSethas no way to say "here's the exact prompt asset (and version) that produced this test case's output." Wiring the two together would give a real input → output → grade trace without either project owning the other's format — e.g. an optional field in the frontmatter (or intools/pipeline metadata) referencing an EvalPort suite id, or on our side, aResultSet.metadataconvention for referencing a.promptyname + version.I also noticed Prompty already has a
GuardrailResultcontract (allowed/reason/rewrite) inschema/model/contracts/guardrails/— a different shape than EvalPort's grader results (ours carry pass/fail + score + explanation per test case), but conceptually adjacent, which is part of what made me think this was worth raising rather than assuming a fit.Not proposing any code right now — just flagging the shape of the connection in case it's useful, since Prompty is clearly already thinking about this space (the
llm-evaluationtopic tag, the.tracytrace viewer). EvalPort's spec is here if useful context: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.mdNo pressure at all if this isn't a priority right now, especially with v2 still in alpha — happy to leave this here for whenever it's useful, or hear that it isn't a fit.
Thanks for Prompty — "write the prompt once, run it from VS Code, Python, or TypeScript" is a nice piece of infrastructure.
All reactions