unilab.envs.manager_based_rl_env.ManagerBasedRlEnv¶
- class unilab.envs.manager_based_rl_env.ManagerBasedRlEnv[source]¶
Bases:
NpEnvManager-Based API adapter that reuses the single
NpEnvlifecycle.- Parameters:
cfg (
ManagerBasedRlEnvCfg)backend (
SimBackend)num_envs (
int)
Methods
__init__(cfg, backend, num_envs)apply_action(actions, state)Subclasses implement the action-to-control conversion.
Capture one detached RGB video frame through the env contract.
close()Close the environment and release backend-owned scene assets.
Export cumulative training progress, independently of episode/physics state.
Return a detached physics snapshot for offline playback/video export.
Aggregate playback overlays from command terms that provide one.
get_playback_model([env_index])Return the backend playback model for one env in a vectorized batch.
Return the backend scene visual model file on the cold path, when available.
import_training_state(state)Restore the authoritative counter and its manager-derived counters.
init_play_renderer([render_spacing, ...])Initialize backend-native playback rendering when available.
Initialize environment and return initial state
render([mode])Render the current state to an RGB array through the play renderer.
Render one interactive playback frame through the env contract.
reset([env_indices, seed, env_ids, options])resolve_play_render_plan(*, ...)Resolve high-level playback mode through the concrete backend.
run_playback(*, initialize, step, num_steps)Execute playback through the concrete backend.
run_playback_mode(*, play_render_mode, ...)Resolve configured playback mode and execute it through the backend contract.
seed([seed])set_autoreset(enabled)Toggle automatic reset of done envs at the end of
step.set_episode_length_buf(values)Overwrite per-env episode counters (cold path, runner-init only).
set_nan_guard(guard)step(actions)Step the environment with given actions, return new state
update_state(state)Subclasses compute observation, reward, and termination state.
Attributes
Action space
The configuration of the environment
return the size of the env if it is vectorized
101}.
Observation space
Return env-facing play/render capabilities.
Current environment state (None before first reset)
- is_vector_env = True¶
- event_manager: EventManager¶
- command_manager: CommandManager | NullCommandManager¶
- action_manager: ActionManager¶
- observation_manager: ObservationManager¶
- termination_manager: TerminationManager¶
- reward_manager: RewardManager¶
- curriculum_manager: CurriculumManager | NullCurriculumManager¶
- metrics_manager: MetricsManager | NullMetricsManager¶
- recorder_manager: RecorderManager | NullRecorderManager¶
- __init__(cfg, backend, num_envs)[source]¶
- Parameters:
cfg (
ManagerBasedRlEnvCfg)backend (
SimBackend)num_envs (
int)
- obs_buf: dict[str, np.ndarray]¶
- extras: dict[str, Any]¶
- property obs_groups_spec: dict[str, int]¶
101}.
Subclasses MUST override this property.
- Type:
Return observation group dimensions, e.g. {“obs”
- Type:
98, “critic”
- property unwrapped: ManagerBasedRlEnv¶
- get_playback_debug_overlays()[source]¶
Aggregate playback overlays from command terms that provide one.
Terms opt in by implementing
playback_debug_overlay_getter(), which returns a per-frame getter following theunisim.backend.base.DebugOverlayGettercontract. The returned getter merges primitives per env across all providing terms. ReturnsNonewhen no term provides an overlay.- Return type:
DebugOverlayGetter | None
- step(actions)[source]¶
Step the environment with given actions, return new state
- Parameters:
actions (
ndarray)- Return type:
- apply_action(actions, state)[source]¶
Subclasses implement the action-to-control conversion.
- Parameters:
actions (
ndarray)state (
NpEnvState)
- Return type:
- update_state(state)[source]¶
Subclasses compute observation, reward, and termination state.
- Parameters:
state (
NpEnvState)- Return type:
- set_episode_length_buf(values)[source]¶
Overwrite per-env episode counters (cold path, runner-init only).
episode_length_bufmirrorsstate.info["steps"]: each step writesinfo["steps"] + 1into the buffer andreset()zeroes both, so a direct assignment must update the two together. RL runners (e.g. RSL-RLinit_at_random_ep_len) call this once before learning starts to stagger initial episode lengths across envs.
- import_training_state(state)[source]¶
Restore the authoritative counter and its manager-derived counters.
- capture_play_video_frame()¶
Capture one detached RGB video frame through the env contract.
- Return type:
- export_training_state()¶
Export cumulative training progress, independently of episode/physics state.
Task-specific curriculum state belongs to the task’s explicit provider; this payload deliberately does not inspect manager or environment internals.
- get_physics_state_snapshot()¶
Return a detached physics snapshot for offline playback/video export.
- Return type:
- get_playback_model(env_index=None)¶
Return the backend playback model for one env in a vectorized batch.
- get_scene_visual_model_file()¶
Return the backend scene visual model file on the cold path, when available.
- init_play_renderer(render_spacing=None, render_offset_mode=None, *, headless=False, capture=False, width=1280, height=720, camera_kwargs=None)¶
Initialize backend-native playback rendering when available.
- property play_capabilities: EnvPlayCapabilities¶
Return env-facing play/render capabilities.
- render(mode='rgb_array')¶
Render the current state to an RGB array through the play renderer.
Lazily initializes a headless capture renderer on first use and returns one detached
(H, W, 3)uint8 frame per call. Backends without native video capture fail closed with a class-named error.
- render_play_frame()¶
Render one interactive playback frame through the env contract.
- Return type:
- resolve_play_render_plan(*, play_render_mode, play_steps, output_video)¶
Resolve high-level playback mode through the concrete backend.
- run_playback(*, initialize, step, num_steps, output_video=None, render_spacing=None, render_offset_mode=None, headless=None, record_video=None, frame_state_getter=None, camera_kwargs=None, debug_overlay_getter=None, on_frame=None)¶
Execute playback through the concrete backend.
on_frameis declared on the env contract but the unisimSimBackend.run_playbackboundary does not accept it yet; passing a callback fails closed until the upstream contract lands. Useunilab.visualization.playback_session.SnapshotPlaybackSessionfor deferred rendering with per-frame callbacks today.- Parameters:
initialize (Callable[[], Any])
step (Callable[[Any], Any])
num_steps (int | None)
output_video (str | PathLike[str] | None)
render_spacing (float | None)
render_offset_mode (str | None)
headless (bool | None)
record_video (bool | None)
frame_state_getter (Callable[[], np.ndarray] | None)
camera_kwargs (CameraCfg | Mapping[str, Any] | None)
debug_overlay_getter (DebugOverlayGetter | None)
on_frame (Callable[[int, np.ndarray], np.ndarray | None] | None)
- Return type:
str | None
- run_playback_mode(*, play_render_mode, play_steps, output_video, initialize, step, render_spacing=None, render_offset_mode=None, frame_state_getter=None, camera_kwargs=None, debug_overlay_getter=None, on_frame=None, on_plan=None)¶
Resolve configured playback mode and execute it through the backend contract.
Overlay gating follows the resolved plan mode:
recordplayback requiresplay_capabilities.supports_debug_overlay(backends fail closed otherwise);interactiveplayback requiressupports_interactive_debug_overlay— when the backend lacks it the getter is dropped with a warning so interactive playback still runs without overlays.
- set_autoreset(enabled)¶
Toggle automatic reset of done envs at the end of
step.Defaults to
True(standard RL autoreset). Interactive playback can disable it so a terminated robot stays put until a manual reset.
- property state: NpEnvState | None¶
Current environment state (None before first reset)