Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions agent_context/topics/env-framework/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,8 @@ inline `robot`, `sensor`, and scene fields continue to parse unchanged.
`build_env_cfg_from_args()` expands `environment.component` before applying
launcher arguments so environment-owned run controls such as `max_episodes`
remain visible to the run loop while explicit CLI values retain precedence.
New offline collection configs should use `collection.target_episodes`; the
legacy field and `--max_episodes` option map to that committed-row target.

An inline runnable config must declare exactly one
`physics: default|newton` backend. A reusable environment component also owns
Expand Down
18 changes: 9 additions & 9 deletions docs/architecture/curated.json
Original file line number Diff line number Diff line change
Expand Up @@ -976,8 +976,8 @@
{
"path": "embodichain/lab/gym/utils/gym_utils.py",
"symbol": "build_env_cfg_from_args",
"start_line": 1210,
"end_line": 1210,
"start_line": 1457,
"end_line": 1457,
"excerpt": "def build_env_cfg_from_args("
}
],
Expand Down Expand Up @@ -3819,8 +3819,8 @@
{
"path": "embodichain/lab/scripts/run_env.py",
"symbol": "cli",
"start_line": 1023,
"end_line": 1023,
"start_line": 1628,
"end_line": 1628,
"excerpt": " env_cfg, gym_config, action_config = build_env_cfg_from_args(args)"
}
]
Expand Down Expand Up @@ -3862,15 +3862,15 @@
{
"path": "embodichain/lab/gym/utils/gym_utils.py",
"symbol": "build_env_cfg_from_args",
"start_line": 1244,
"end_line": 1249,
"excerpt": " cfg: EmbodiedEnvCfg = config_to_cfg(\n gym_config,\n manager_modules=get_manager_modules(),\n source_path=gym_config_source_path,\n task_program_path_override=getattr(args, \"task_program\", None),\n )"
"start_line": 1497,
"end_line": 1502,
"excerpt": " cfg: EmbodiedEnvCfg = config_to_cfg(\n gym_config,\n manager_modules=get_manager_modules(),\n source_path=gym_config_source_path,\n task_program_path_override=getattr(args, \"task_program\", None),\n )"
Comment on lines +3865 to +3867

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Evidence no longer matches revision

The snapshot is pinned to revision 3224ac1, but this excerpt and the other updated anchors describe the PR’s current source. Direct validation compares them with the pinned revision and reports stale evidence. The architecture links also no longer point to the code the snapshot claims to document.

Prompt To Fix With AI
This is a comment left during a code review.
Path: docs/architecture/curated.json
Line: 3865-3867

Comment:
**Evidence no longer matches revision**

The snapshot is pinned to revision `3224ac1`, but this excerpt and the other updated anchors describe the PR’s current source. Direct validation compares them with the pinned revision and reports stale evidence. The architecture links also no longer point to the code the snapshot claims to document.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Codex Fix in Claude Code

},
{
"path": "embodichain/lab/scripts/run_env.py",
"symbol": "cli",
"start_line": 1028,
"end_line": 1028,
"start_line": 1680,
"end_line": 1680,
"excerpt": " env = gymnasium.make(id=gym_config[\"id\"], cfg=env_cfg, **action_config)"
}
]
Expand Down
30 changes: 18 additions & 12 deletions docs/source/guides/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -129,10 +129,11 @@ For RL training and data expansion, EmbodiChain uses file-based configs (`.json`

Configs are loaded with `embodichain.utils.utility.load_config`, which selects the parser from the file extension. Both formats produce the same in-memory dictionary and are passed to `config_to_cfg()` for environment setup.

For offline expert expansion, `max_episodes` counts persisted
per-environment episodes rather than vector batches. Thus `num_envs: 4` and
`max_episodes: 10` produce two full four-row commits plus a final two-row
commit. Failed rows count only when the relevant `DatasetFunctorCfg` sets
For offline expert expansion, `collection.target_episodes` counts persisted
environment rows rather than vector batches. Thus `num_envs: 4` and
`target_episodes: 10` produce two full four-row commits plus a final two-row
commit. The legacy `max_episodes` field and CLI option map to the same target.
Failed rows count only when the relevant `DatasetFunctorCfg` sets
`save_failed_episodes: true`.

Example paths in the repository:
Expand All @@ -150,7 +151,9 @@ When a training config references a gym config (via `trainer.gym_config`), the n
{
"id": "EmbodiedEnv-v1",
"num_envs": 4,
"max_episodes": 100,
"collection": {
"target_episodes": 100
},
"max_episode_steps": 600,
"physics": "default",
"device": "cpu",
Expand Down Expand Up @@ -392,13 +395,16 @@ canonical `entity_id` to physical `simulation_uid` mapping.

The optional `expansion.config` reference keeps task-facing augmentation
settings in a separate file. That file may define `runtime`, `policy`,
`overrides`, and `candidate_indices`; the runner deep-merges task-local fields,
applies the runtime overlay after resolving the selected environment variant,
and then constructs the common environment. `--physics` selects `default` or
`newton` from the task mapping, while `--expansion-profile` and
`--expansion-candidate-indices` override the referenced expansion values.
Relative policy and resource paths inside the expansion file resolve from that
file's directory.
`overrides`, and `collection`; `collection.target_episodes` is the final
committed row count, while `collection.selection` chooses sequential or
explicit logical recipes. The runner deep-merges task-local fields, applies
the runtime overlay after resolving the selected environment variant, and then
constructs the common environment. `--physics` selects `default` or `newton`
from the task mapping, while `--expansion-profile` and
`--expansion-recipe-indices` override the referenced expansion values. The
legacy `candidate_indices` spelling remains accepted as an explicit recipe
selection alias. Relative policy and resource paths inside the expansion file
resolve from that file's directory.

Component ownership is exclusive. Do not combine `environment.component` with
inline environment or scene fields, and do not combine `embodiment.component`
Expand Down
84 changes: 70 additions & 14 deletions docs/source/guides/run_env.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,11 +77,66 @@ expansion:
```

The referenced file may contain `runtime`, `policy`, `overrides`, and
`candidate_indices`. The runner resolves it relative to the task, applies its
runtime overlay after expanding the selected environment variant, and merges
task-local expansion fields on top. `run-task` then uses that declaration for
Task Program expansion; `--expansion-profile` and
`--expansion-candidate-indices` remain per-run overrides.
`collection`. The runner resolves it relative to the task, applies its runtime
overlay after expanding the selected environment variant, and merges task-local
expansion fields on top. `run-task` then uses that declaration for Task Program
expansion. `--expansion-profile` remains a per-run profile override;
`--expansion-recipe-indices` selects explicit logical recipes. The legacy
`--expansion-candidate-indices` spelling is still accepted as an alias.

### Episode targets and batches

All offline collection paths use one target count:

```yaml
collection:
target_episodes: 64
max_attempts: 3
selection:
mode: sequential
start_recipe_index: 0
```

`target_episodes` is the number of successfully committed environment rows.
`num_envs` only sets the maximum width of one parallel batch, so
`target_episodes: 64` with `num_envs: 16` runs four batches. A final partial
batch uses only the rows needed to reach the target. `max_attempts` applies to
one selected batch attempt and failed attempts are discarded before retrying.

The legacy `--max_episodes` CLI option remains supported and maps to
`collection.target_episodes`. An environment component's `max_episodes` is the
fallback when no collection target is declared. A task-level collection target
takes precedence over that component default; conflicting explicit target
values are rejected during parsing.

Expansion selection has two modes:

```yaml
# Select recipe 0 through recipe 63 and split them into batches automatically.
collection:
target_episodes: 64
selection:
mode: sequential
start_recipe_index: 0

# Select exactly four logical recipes and collect four rows.
collection:
target_episodes: 4
selection:
mode: explicit
recipe_indices: [0, 16, 32, 48]
```

Recipe indices identify logical expansion candidates. They do not identify
batch starts. Affordance branches, scene families, trajectory variants, and
visual profiles change the selected episode's content; they do not increase
the target count.

Collection manifests record `target_episodes`, `planned_episodes`,
`committed_episodes`, `rejected_episodes`, `attempts`, `batch_count`, and the
prepare, commit, and discard reset counts. A prepare reset happens before a
batch, a commit reset finalizes selected rows, and a discard reset clears a
failed attempt. None of these reset boundaries adds an episode.

The runnable config, or its selected environment component, must declare
`physics: default` or `physics: newton`. With an `environment.default/newton`
Expand Down Expand Up @@ -162,11 +217,11 @@ details.

Without `--preview` or `--replay`, `run-env` enters offline rollout mode. For
each vector batch, it asks the task for its demonstration segments, applies
every action through `env.step()`, and commits selected environment rows with
an explicit reset. `max_episodes` is the exact number of persisted
per-environment episodes, not the number of vector batches. For example,
`max_episodes=10` with `num_envs=4` runs three batches and commits only two rows
from the final batch. `--max_episodes` overrides the value in the gym config:
every action through `env.step()`, and commits selected environment rows with an explicit reset.
`collection.target_episodes` is the exact number of persisted environment rows,
not the number of vector batches. For example, `target_episodes=10` with
`num_envs=4` runs three batches and commits only two rows from the final batch.
The legacy `--max_episodes` option maps to the same target:

```bash
embodichain run-env \
Expand All @@ -182,9 +237,10 @@ Headless execution is normally preferred for throughput. Use
structured dataset.

Failed attempts are discarded and retried by default, up to
`demo_max_attempts` (default: 3). Set `save_failed_episodes: true` on a dataset
functor to keep a failed or truncated attempt that contains recorded frames.
Such a commit counts toward `max_episodes` and is not retried. Empty plans and
`collection.max_attempts` (with `demo_max_attempts` retained as a legacy
fallback). Set `save_failed_episodes: true` on a dataset functor to keep a
failed or truncated attempt that contains recorded frames. Such a commit
counts toward `collection.target_episodes` and is not retried. Empty plans and
exceptions have no complete dataset transaction and are still discarded.

### Multi-segment episodes
Expand Down Expand Up @@ -253,7 +309,7 @@ continue. Consequently, rollout and trajectory lengths may differ by row.
The executor's result remains batch-atomic: without failed-data saving every
row must eventually succeed, while any failure or truncation invalidates the
batch. With `save_failed_episodes`, selected failed rows are committed with
their per-row failure metadata. Rows not needed to reach `max_episodes` are
their per-row failure metadata. Rows not needed to reach `collection.target_episodes` are
explicitly discarded, so parallel collection never overshoots the requested
episode count.

Expand Down
5 changes: 3 additions & 2 deletions docs/source/overview/gym/dataset_functors.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,8 +76,9 @@ is ``false``. During ``run-env`` expert generation:
row contains at least one frame. The saved sidecar records ``success=false``
and the terminal reason;
- empty plans and exceptions are always discarded; and
- a saved failure counts toward ``max_episodes``, so the requested dataset size
is not exceeded.
- a saved failure counts toward ``collection.target_episodes`` (with legacy
``max_episodes`` mapped to that target), so the requested dataset size is
not exceeded.

When several save-mode functors are configured, enabling this option on any of
them makes the Dataset Manager submit failed rows to all save-mode functors so
Expand Down
14 changes: 11 additions & 3 deletions docs/source/overview/task_program/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -190,7 +190,13 @@ simulation:
env:
sim_steps_per_control: 4
events: {}
dataset: {}
dataset:
lerobot:
func: LeRobotRecorder
mode: save
params:
save_path: /tmp/repeated-pick-place/datasets
# robot_type is filled from the selected embodiment.
```

A runnable `task.<embodiment>.yaml` selects the environment variants, all three
Expand All @@ -214,8 +220,10 @@ expansion:

The referenced expansion file is task-facing configuration rather than a
second environment deployment. It owns batch runtime values, the shared
expansion policy, task-specific augmentation overrides, and candidate indices;
the runner merges it after selecting the physical backend.
expansion policy, task-specific augmentation overrides, and the collection
target and recipe selection; the runner merges it after selecting the physical
backend. `collection.target_episodes` is the final committed row count, while
`num_envs` only controls the width of each batch.

The callable-free `integration.yaml` owns the semantic scene binding and
task-specific profile additions. Its canonical identities map explicitly to
Expand Down
80 changes: 72 additions & 8 deletions docs/source/tutorial/data_expansion.rst
Original file line number Diff line number Diff line change
Expand Up @@ -333,14 +333,15 @@ functor under ``env.dataset`` in an inline Gym config or in a reusable

Important fields are:

* ``max_episodes`` is the exact number of persisted per-environment episodes,
not the number of vector batches.
* ``collection.target_episodes`` is the exact number of persisted environment
rows, not the number of vector batches. The legacy ``max_episodes`` setting
and ``--max_episodes`` option map to this target.
* ``max_episode_steps`` must exceed the longest valid expert execution,
including gripper holds and settling actions.
* ``save_failed_episodes`` belongs beside ``func`` and ``mode``. It defaults to
``false``; when enabled, a failed or truncated attempt with recorded frames
is committed with ``success=false`` metadata and counts toward
``max_episodes``.
``collection.target_episodes``.
* ``params.save_path`` is the parent directory for auto-numbered datasets. If
omitted, the default is ``~/.cache/embodichain_datasets`` or the value of
``EMBODICHAIN_DATASET_ROOT``.
Expand Down Expand Up @@ -373,14 +374,76 @@ data expansion:
attempt with ``reset(options={"save_data": False})``.

Failed attempts are discarded and retried by default, up to
``demo_max_attempts`` (default: 3). Empty plans and exceptions are always
``collection.max_attempts`` (with ``demo_max_attempts`` as a legacy fallback).
Empty plans and exceptions are always
discarded because they do not form a complete dataset transaction. With
``save_failed_episodes: true``, a failed or truncated attempt is retained only
when every selected row contains recorded frames.

``num_envs`` controls collection parallelism. If ``max_episodes=10`` and
``num_envs=4``, the runner uses three vector batches and commits only two rows
from the final batch, so it never overshoots the requested episode count.
``collection.target_episodes`` controls the final dataset size and ``num_envs``
controls only the parallel width. If ``target_episodes=10`` and ``num_envs=4``,
the runner uses three vector batches and commits only two rows from the final
batch. The legacy ``--max_episodes`` option maps to the same target field.

For a configured Expansion task, keep the collection policy beside the
task-facing expansion overrides:

.. code-block:: yaml

runtime:
num_envs: 16
max_episode_steps: 1200
collection:
target_episodes: 64
max_attempts: 3
selection:
mode: sequential
start_recipe_index: 0

Sequential selection consumes logical recipe IDs from the start index and
automatically groups them into batches. Explicit selection is useful when a
small, known set of recipes is required:

.. code-block:: yaml

collection:
target_episodes: 4
selection:
mode: explicit
recipe_indices: [0, 16, 32, 48]

The explicit recipe list must contain exactly ``target_episodes`` entries.
Recipe IDs identify logical candidates; they do not represent batch offsets.

Collection terms have one meaning across handwritten and Expansion tasks:

.. list-table:: Collection terms
:header-rows: 1
:widths: 20 55

* - Term
- Meaning
* - ``episode``
- One environment row executed to completion and submitted as data.
* - ``attempt``
- One execution try. A failed attempt can be discarded and retried.
* - ``batch``
- The selected environment rows executed in one vectorized pass.
* - ``num_envs``
- The maximum number of rows available in one batch.
* - ``recipe_index``
- The stable logical ID of an Expansion candidate.

Changing ``num_envs`` changes the number of batches and the parallel capacity;
it does not change ``target_episodes``. A failed attempt and any prepare,
commit, or discard reset also leave the target episode count unchanged.

Collection accounting distinguishes successful data rows from execution
attempts. The expansion manifest records the target, planned, committed, and
rejected episode counts, total attempts, batch count, and prepare, commit, and
discard reset counts. Prepare reset runs before a batch, commit reset finalizes
selected rows, and discard reset clears a failed attempt. Reset boundaries do
not add episodes.

An episode is the complete task; a segment is one semantic subtask. Do not use
``generate_function(num_traj=...)`` to repeat subtasks: direct callers may pass
Expand All @@ -394,7 +457,8 @@ Useful modes and options are:
* ``--filter_dataset_saving`` executes the expert while suppressing structured
dataset writes.
* ``--num_envs`` overrides collection parallelism.
* ``--max_episodes`` overrides the configured episode target.
* ``--max_episodes`` overrides the configured ``collection.target_episodes``.
* ``collection.max_attempts`` controls retries for one selected batch.

See :doc:`/guides/run_env` for preview, dataset recording, debug video,
trajectory recording, and replay modes, and :doc:`/guides/cli` for the complete
Expand Down
Loading
Loading