Skip to content

dataset(code-review): add synthetic reaction gold answers from BCApps PR 11063 - #863

Draft
Wenjie Fan (gggdttt) wants to merge 1 commit into
mainfrom
self-improvement/bcapps-11063-synthetic-eval
Draft

dataset(code-review): add synthetic reaction gold answers from BCApps PR 11063#863
Wenjie Fan (gggdttt) wants to merge 1 commit into
mainfrom
self-improvement/bcapps-11063-synthetic-eval

Conversation

@gggdttt

@gggdttt Wenjie Fan (gggdttt) commented Sep 8, 2026

Copy link
Copy Markdown
Collaborator

Summary

Adds minimal synthetic code-review gold answers from explicit BCApps PR feedback.

Full source PR patches are intentionally excluded. Positive reactions become
isolated expected findings; thumbs-down findings become precision guards by
omission. Existing equivalent gold answers are not duplicated.

Accepted findings

Rejected findings

Knowledge dependency

Validation

Result: NOT supported by the evaluation - this pull request is not ready to merge.

The run emitted 2 unexpected finding(s) and missed 0 expected finding(s).

Revision under test

Component Revision
BC-Bench evaluation branch 41b21139b86542105aabf572a1cb44d2c54e8961 (self-improvement/bcapps-11063-synthetic-eval-eval)
BC-ALAgents engine 159572aad814d6e1022ca8d54e85a6eab7c4d5c6
BCQuality knowledge e74e2a6601ec87a1ffec6e3d7eb074aa67437ef2

The engine commit pins the BCQuality commit, and the evaluation branch pins the engine commit, so the revision the run executed equals the revision it declares. No run-time override was used.

Run: https://github.com/microsoft/BC-Bench/actions/runs/34203457236 (conclusion success, wall clock 7.1 min)

Per entry

Entry F1
synthetic__perf-watermark-folder-sync-01 0

Aggregate

Metric Value
total 1
expected_comment_count 0
generated_comment_count 2
matched_comment_count 0
incorrect_comment_count 2
missed_comment_count 0
ignored_matched_comment_count 0
precision 0
recall 1
f1 0

Cost: n/a average AI credits per entry, 232.3s average duration.
Benchmark 0.11.0, model gpt-5-6-luna, judge gpt-5.3-codex.

Schema checks: loaded the complete dataset/codereview.jsonl corpus through CodeReviewEntry and ran the targeted code-review and dataset-integrity tests.

Generated by the BC-ALAgentsInternal self-improvement workflow.

Offline evaluation: regression

Candidate correctness failed: unexpected or missed findings remain. Existing ignored gold comments retain their neutral scoring semantics.

Common engine code: ecf8e31759d6ddd6d78e3a0b7836b40134368009; dataset SHA256: ACDC50C36CC88DD31A5AB628A7C3D06A44E5F7628DA9DA79A0916E6AF0F46222.
Model: gpt-5.6-luna; judge: gpt-5.3-codex. Exact D: synthetic__perf-watermark-folder-sync-01.
Dataset base: 231538937b19317eddc3ac654459ca29b3ef4170; candidate: 46b6ad94497fcb272417b170b2cbb7b3f96d7209.

Arm Evaluation commit Engine pin Knowledge pin
baseline 27da12f4d5dfec9862f23219b9d1e80c465cc18a 4407f20b44fe92e3693d0444d25cc3430f1d5262 8584217c7506eea7eef27a9db353d78c0cd6a9a1
candidate 09695a2fecf379e823762ccbc425a09b9259a823 35b7bcdecf431b9b1a18ca83830a38c2fc353344 e74e2a6601ec87a1ffec6e3d7eb074aa67437ef2

Baseline

Run: https://github.com/microsoft/BC-Bench/actions/runs/34221111997; conclusion: success; wall clock: 12.6 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.

Entry Expected Generated Missed Unexpected F1 AI credits Agent seconds
synthetic__perf-watermark-folder-sync-01 0 2 0 2 0 unavailable 572.746033486

Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total.

Aggregate metric Value
total 1
expected_comment_count 0
generated_comment_count 2
matched_comment_count 0
missed_comment_count 0
incorrect_comment_count 2
precision 0
recall 1
f1 0

Candidate

Run: https://github.com/microsoft/BC-Bench/actions/runs/34222250479; conclusion: success; wall clock: 7 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.

Entry Expected Generated Missed Unexpected F1 AI credits Agent seconds
synthetic__perf-watermark-folder-sync-01 0 1 0 1 0 unavailable 234.327608426

Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total.

Aggregate metric Value
total 1
expected_comment_count 0
generated_comment_count 1
matched_comment_count 0
missed_comment_count 0
incorrect_comment_count 1
precision 0
recall 1
f1 0

Missing telemetry is unavailable, not zero. Evaluation-only metrics exclude candidate generation and are not the full-cycle cost.
Human review must verify source-patch fidelity, gold correctness, and target attribution. No automatic merge or branch-protection claim is made.

@gggdttt
Wenjie Fan (gggdttt) marked this pull request as draft September 8, 2026 08:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant