VIPRewardTransform¶

class torchrl.envs.transforms.VIPRewardTransform(*args, **kwargs)[來源]¶

一個 VIP 變換，用於根據嵌入相似性計算獎勵。

此類將更新獎勵計算

forward(tensordict)[來源]¶

讀取輸入 tensordict，並對選定的鍵應用轉換。

預設情況下，此方法

直接呼叫 _apply_transform()。
不呼叫 _step() 或 _call()。

此方法不會在任何時候在 env.step 中呼叫。但是，它會在 sample() 中呼叫。

注意

forward 也可以使用 dispatch 將引數名稱轉換為鍵，並使用常規關鍵字引數。

示例

>>> class TransformThatMeasuresBytes(Transform):
...     '''Measures the number of bytes in the tensordict, and writes it under `"bytes"`.'''
...     def __init__(self):
...         super().__init__(in_keys=[], out_keys=["bytes"])
...
...     def forward(self, tensordict: TensorDictBase) -> TensorDictBase:
...         bytes_in_td = tensordict.bytes()
...         tensordict["bytes"] = bytes
...         return tensordict
>>> t = TransformThatMeasuresBytes()
>>> env = env.append_transform(t) # works within envs
>>> t(TensorDict(a=0))  # Works offline too.

transform_input_spec(input_spec: TensorSpec) → TensorSpec[來源]¶

轉換輸入規範，使結果規範與轉換對映匹配。

引數:: input_spec (TensorSpec) – 轉換前的規範
返回:: 轉換後的預期規範

VIPRewardTransform¶

文件

教程

資源