fovi.sensing.projection

Calibrated central cameras and spherical gaze, without renderer dependencies.

Pixels use OpenCV’s integer pixel-center convention. Directions use camera coordinates X right, Y down, Z forward. External adapters convert frames once.

class fovi.sensing.projection.Mapping

Bases: Collection

A Mapping is a generic container for associating key/value pairs.

This class provides concrete generic implementations of all methods except for __getitem__, __iter__, and __len__.

get(k[, d]) → D[k] if k in D, else d.  d defaults to None.
items() → a set-like object providing a view on D's items
keys() → a set-like object providing a view on D's keys
values() → an object providing a view on D's values
fovi.sensing.projection.AbstractSet

alias of Set

fovi.sensing.projection.dataclass(cls=None, /, *, init=True, repr=True, eq=True, order=False, unsafe_hash=False, frozen=False, match_args=True, kw_only=False, slots=False, weakref_slot=False)[source]

Add dunder methods based on the fields defined in the class.

Examines PEP 526 __annotations__ to determine fields.

If init is true, an __init__() method is added to the class. If repr is true, a __repr__() method is added. If order is true, rich comparison dunder methods are added. If unsafe_hash is true, a __hash__() method is added. If frozen is true, fields may not be assigned to after instance creation. If match_args is true, the __match_args__ tuple is added. If kw_only is true, then by default all fields are keyword-only. If slots is true, a new class with a __slots__ attribute is returned.

fovi.sensing.projection.replace(obj, /, **changes)[source]

Return a new object replacing specified fields with new values.

This is especially useful for frozen classes. Example usage:

@dataclass(frozen=True)
class C:
    x: int
    y: int

c = C(1, 2)
c1 = replace(c, x=3)
assert c1.x == 3 and c1.y == 2
class fovi.sensing.projection.Real[source]

Bases: Complex

To Complex, Real adds the operations that work on real numbers.

In short, those are: a conversion to float, trunc(), divmod, %, <, <=, >, and >=.

Real also provides defaults for the derived operations.

property real

Real numbers are their real component.

property imag

Real numbers have no imaginary component.

conjugate()[source]

Conjugate is a no-op for Reals.

fovi.sensing.projection.TypedDict(typename, fields=None, /, *, total=True, **kwargs)[source]

A simple typed namespace. At runtime it is equivalent to a plain dict.

TypedDict creates a dictionary type such that a type checker will expect all instances to have a certain set of keys, where each key is associated with a value of a consistent type. This expectation is not checked at runtime.

Usage:

>>> class Point2D(TypedDict):
...     x: int
...     y: int
...     label: str
...
>>> a: Point2D = {'x': 1, 'y': 2, 'label': 'good'}  # OK
>>> b: Point2D = {'z': 3, 'label': 'bad'}           # Fails type check
>>> Point2D(x=1, y=2, label='first') == dict(x=1, y=2, label='first')
True

The type info can be accessed via the Point2D.__annotations__ dict, and the Point2D.__required_keys__ and Point2D.__optional_keys__ frozensets. TypedDict supports an additional equivalent form:

Point2D = TypedDict('Point2D', {'x': int, 'y': int, 'label': str})

By default, all keys must be present in a TypedDict. It is possible to override this by specifying totality:

class Point2D(TypedDict, total=False):
    x: int
    y: int

This means that a Point2D TypedDict can have any of the keys omitted. A type checker is only expected to support a literal False or True as the value of the total argument. True is the default, and makes all items defined in the class body be required.

The Required and NotRequired special forms can also be used to mark individual keys as being required or not required:

class Point2D(TypedDict):
    x: int               # the "x" key must always be present (Required is the default)
    y: NotRequired[int]  # the "y" key can be omitted

See PEP 655 for more details on Required and NotRequired.

class fovi.sensing.projection.Tensor

Bases: TensorBase

_clear_non_serializable_cached_data()[source]

Clears any data cached in the tensor’s __dict__ that would prevent the tensor from being serialized.

For example, subclasses with custom dispatched sizes / strides cache this info in non-serializable PyCapsules within the __dict__, and this must be cleared out for serialization to function.

Any subclass that overrides this MUST call super()._clear_non_serializable_cached_data(). Additional data cleared within the override must be able to be re-cached transparently to avoid breaking subclass functionality.

backward(gradient=None, retain_graph=None, create_graph=False, inputs=None)[source]

Computes the gradient of current tensor wrt graph leaves.

The graph is differentiated using the chain rule. If the tensor is non-scalar (i.e. its data has more than one element) and requires gradient, the function additionally requires specifying a gradient. It should be a tensor of matching type and shape, that represents the gradient of the differentiated function w.r.t. self.

This function accumulates gradients in the leaves - you might need to zero .grad attributes or set them to None before calling it. See Default gradient layouts for details on the memory layout of accumulated gradients.

Note

If you run any forward ops, create gradient, and/or call backward in a user-specified CUDA stream context, see Stream semantics of backward passes.

Note

When inputs are provided and a given input is not a leaf, the current implementation will call its grad_fn (though it is not strictly needed to get this gradients). It is an implementation detail on which the user should not rely. See https://github.com/pytorch/pytorch/pull/60521#issuecomment-867061780 for more details.

Parameters:
  • gradient (Tensor, optional) – The gradient of the function being differentiated w.r.t. self. This argument can be omitted if self is a scalar. Defaults to None.

  • retain_graph (bool, optional) – If False, the graph used to compute the grads will be freed; If True, it will be retained. The default is None, in which case the value is inferred from create_graph (i.e., the graph is retained only when higher-order derivative tracking is requested). Note that in nearly all cases setting this option to True is not needed and often can be worked around in a much more efficient way.

  • create_graph (bool, optional) – If True, graph of the derivative will be constructed, allowing to compute higher order derivative products. Defaults to False.

  • inputs (Sequence[Tensor] or dict[str, Tensor], optional) – Inputs w.r.t. which the gradient will be accumulated into .grad. All other tensors will be ignored. If not provided, the gradient is accumulated into all the leaf Tensors that were used to compute the tensors. A dict of tensors (e.g. dict(model.named_parameters())) is also accepted. Defaults to None.

cholesky(upper=False)[source]
detach()

Returns a new Tensor, detached from the current graph.

The result will never require gradient.

This method also affects forward mode AD gradients and the result will never have forward mode AD gradients.

Note

Returned Tensor shares the same storage with the original one. In-place modifications on either of them will be seen, and may trigger errors in correctness checks.

detach_()

Detaches the Tensor from the graph that created it, making it a leaf. Views cannot be detached in-place.

This method also affects forward mode AD gradients and the result will never have forward mode AD gradients.

dim_order(ambiguity_check=False) → tuple[source]

Returns the uniquely determined tuple of int describing the dim order or physical layout of self.

The dim order represents how dimensions are laid out in memory of dense tensors, starting from the outermost to the innermost dimension.

Note that the dim order may not always be uniquely determined. If ambiguity_check is True, this function raises a RuntimeError when the dim order cannot be uniquely determined; If ambiguity_check is a list of memory formats, this function raises a RuntimeError when tensor can not be interpreted into exactly one of the given memory formats, or it cannot be uniquely determined. If ambiguity_check is False, it will return one of legal dim order(s) without checking its uniqueness. Otherwise, it will raise TypeError.

Parameters:

ambiguity_check (bool or List[torch.memory_format]) – The check method for ambiguity of dim order.

Examples:

>>> torch.empty((2, 3, 5, 7)).dim_order()
(0, 1, 2, 3)
>>> torch.empty((2, 3, 5, 7)).transpose(1, 2).dim_order()
(0, 2, 1, 3)
>>> torch.empty((2, 3, 5, 7), memory_format=torch.channels_last).dim_order()
(0, 2, 3, 1)
>>> torch.empty((1, 2, 3, 4)).dim_order()
(0, 1, 2, 3)
>>> try:
...     torch.empty((1, 2, 3, 4)).dim_order(ambiguity_check=True)
... except RuntimeError as e:
...     print(e)
The tensor does not have unique dim order, or cannot map to exact one of the given memory formats.
>>> torch.empty((1, 2, 3, 4)).dim_order(
...     ambiguity_check=[torch.contiguous_format, torch.channels_last]
... )  # It can be mapped to contiguous format
(0, 1, 2, 3)
>>> try:
...     torch.empty((1, 2, 3, 4)).dim_order(ambiguity_check="ILLEGAL") # type: ignore[arg-type]
... except TypeError as e:
...     print(e)
The ambiguity_check argument must be a bool or a list of memory formats.

Warning

The dim_order tensor API is experimental and subject to change.

eig(eigenvectors=False)[source]
index(positions, dims)[source]

Index a regular tensor by binding specified positions to dims.

This converts a regular tensor to a first-class tensor by binding the specified positional dimensions to Dim objects.

Parameters:
  • positions – Tuple of dimension positions to bind

  • dims – Dim objects or tuple of Dim objects to bind to

Returns:

First-class tensor with specified dimensions bound

is_shared()[source]

Checks if tensor is in shared memory.

This is always True for CUDA tensors.

istft(n_fft: int, hop_length: int | None = None, win_length: int | None = None, window: Tensor | None = None, center: bool = True, normalized: bool = False, onesided: bool | None = None, length: int | None = None, return_complex: bool = False)[source]

See torch.istft()

lstsq(other)[source]
lu(pivot=True, get_infos=False)[source]

See torch.lu()

module_load(other, assign=False)[source]

Defines how to transform other when loading it into self in load_state_dict().

Used when get_swap_module_params_on_conversion() is True.

It is expected that self is a parameter or buffer in an nn.Module and other is the value in the state dictionary with the corresponding key, this method defines how other is remapped before being swapped with self via swap_tensors() in load_state_dict().

Note

This method should always return a new object that is not self or other. For example, the default implementation returns self.copy_(other).detach() if assign is False or other.detach() if assign is True.

Parameters:
  • other (Tensor) – value in state dict with key corresponding to self

  • assign (bool) – the assign argument passed to nn.Module.load_state_dict()

norm(p: float | str | None = 'fro', dim=None, keepdim=False, dtype=None)[source]

See torch.linalg.norm()

qr(some=True)[source]
register_hook(hook)[source]

Registers a backward hook.

The hook will be called every time a gradient with respect to the Tensor is computed. The hook should have the following signature:

hook(grad) -> Tensor or None

The hook should not modify its argument, but it can optionally return a new gradient which will be used in place of grad.

This function returns a handle with a method handle.remove() that removes the hook from the module.

Note

See Backward Hooks execution for more information on how when this hook is executed, and how its execution is ordered relative to other hooks.

Example:

>>> v = torch.tensor([0., 0., 0.], requires_grad=True)
>>> h = v.register_hook(lambda grad: grad * 2)  # double the gradient
>>> v.backward(torch.tensor([1., 2., 3.]))
>>> v.grad

 2
 4
 6
[torch.FloatTensor of size (3,)]

>>> h.remove()  # removes the hook
register_post_accumulate_grad_hook(hook)[source]

Registers a backward hook that runs after grad accumulation.

The hook will be called after all gradients for a tensor have been accumulated, meaning that the .grad field has been updated on that tensor. The post accumulate grad hook is ONLY applicable for leaf tensors (tensors without a .grad_fn field). Registering this hook on a non-leaf tensor will error!

The hook should have the following signature:

hook(param: Tensor) -> None

Note that, unlike other autograd hooks, this hook operates on the tensor that requires grad and not the grad itself. The hook can in-place modify and access its Tensor argument, including its .grad field.

This function returns a handle with a method handle.remove() that removes the hook from the module.

Note

See Backward Hooks execution for more information on how when this hook is executed, and how its execution is ordered relative to other hooks. Since this hook runs during the backward pass, it will run in no_grad mode (unless create_graph is True). You can use torch.enable_grad() to re-enable autograd within the hook if you need it.

Example:

>>> v = torch.tensor([0., 0., 0.], requires_grad=True)
>>> lr = 0.01
>>> # simulate a simple SGD update
>>> h = v.register_post_accumulate_grad_hook(lambda p: p.add_(p.grad, alpha=-lr))
>>> v.backward(torch.tensor([1., 2., 3.]))
>>> v
tensor([-0.0100, -0.0200, -0.0300], requires_grad=True)

>>> h.remove()  # removes the hook
reinforce(reward)[source]
resize(*sizes)[source]
resize_as(tensor)[source]
share_memory_()[source]

Moves the underlying storage to shared memory.

This is a no-op if the underlying storage is already in shared memory and for CUDA tensors. Tensors in shared memory cannot be resized.

See torch.UntypedStorage.share_memory_() for more details.

solve(other)[source]
split(split_size, dim=0)[source]

See torch.split()

stft(n_fft: int, hop_length: int | None = None, win_length: int | None = None, window: Tensor | None = None, center: bool = True, pad_mode: str = 'reflect', normalized: bool = False, onesided: bool | None = None, return_complex: bool | None = None, align_to_window: bool | None = None)[source]

See torch.stft()

Warning

This function changed signature at version 0.4.1. Calling with the previous signature may cause error or return incorrect result.

storage() → torch.TypedStorage[source]

Returns the underlying TypedStorage.

Warning

TypedStorage is deprecated. It will be removed in the future, and UntypedStorage will be the only storage class. To access the UntypedStorage directly, use Tensor.untyped_storage().

storage_type() → type[source]

Returns the type of the underlying storage.

symeig(eigenvectors=False)[source]
to_sparse_coo()[source]

Convert a tensor to coordinate format.

Examples:

>>> dense = torch.randn(5, 5)
>>> sparse = dense.to_sparse_coo()
>>> sparse._nnz()
25
unflatten(dim, sizes) → Tensor[source]

See torch.unflatten().

unique(sorted=True, return_inverse=False, return_counts=False, dim=None)[source]

Returns the unique elements of the input tensor.

See torch.unique()

unique_consecutive(return_inverse=False, return_counts=False, dim=None)[source]

Eliminates all but the first element from every consecutive group of equivalent elements.

See torch.unique_consecutive()

fovi.sensing.projection.validate_gaze_convention(convention: str) → None[source]

Reject unsupported gaze conventions.

fovi.sensing.projection._calibration_float(value: Real | Tensor, name: str) → float[source]

Normalize real numeric scalars without accepting strings or containers.

class fovi.sensing.projection.CameraCalibration[source]

Bases: TypedDict

Serializable CameraModel arguments; model, image_size, and intrinsics are required.

model: str
image_size: tuple[int, int]
intrinsics: tuple[float, float, float, float]
distortion: tuple[float, ...]
image_circle: tuple[float, float, float] | None
max_angle_deg: float
class fovi.sensing.projection.CameraModel(model: str, image_size: tuple[int, int], intrinsics: tuple[float, float, float, float], distortion: tuple[float, ...] = (), image_circle: tuple[float, float, float] | None = None, max_angle_deg: float = 90.0)[source]

Bases: object

A calibrated pinhole or equidistant-polynomial fisheye camera.

Parameters:
  • model – pinhole or fisheye (OpenCV fisheye convention).

  • image_size – Source image (height, width).

  • intrinsics – (fx, fy, cx, cy), in integer-center pixel coordinates.

  • distortion – Pinhole (k1, k2, p1, p2[, k3[, k4, k5, k6]]) or fisheye (k1, k2, k3, k4). Empty means the ideal model.

  • image_circle – Optional usable disc (cx, cy, radius), in source pixels.

  • max_angle_deg – Calibrated angular domain about the optical axis.

model: str
image_size: tuple[int, int]
intrinsics: tuple[float, float, float, float]
distortion: tuple[float, ...] = ()
image_circle: tuple[float, float, float] | None = None
max_angle_deg: float = 90.0
classmethod from_config(camera: CameraModel | CameraCalibration) → CameraModel[source]

Normalize serialized calibration once at an API boundary.

resized(image_size: tuple[int, int]) → CameraModel[source]

Calibrate a full-frame resize; cropping and image rotation need new intrinsics.

_pinhole_undistort(target: Tensor) → Tensor[source]

Invert radial/tangential distortion with a batched analytic Newton step.

pixel_validity(pixels: Tensor) → Tensor[source]

Return (…,) validity for (…, 2) source pixels.

project(directions: Tensor) → tuple[Tensor, Tensor][source]

Project (…, 3) directions to (…, 2) pixels and (…,) validity.

unproject(pixels: Tensor) → tuple[Tensor, Tensor][source]

Invert calibrated pixels into unit directions; flag failed inversions.

field_of_view(reference_side: str, fraction: float = 1.0) → float[source]

Return the angular span of a centered short/long-side retinal window.

Window endpoints are image boundaries in integer-center coordinates. The selected axis follows pixel dimensions, not an assumed aspect ratio.

__init__(model: str, image_size: tuple[int, int], intrinsics: tuple[float, float, float, float], distortion: tuple[float, ...] = (), image_circle: tuple[float, float, float] | None = None, max_angle_deg: float = 90.0) → None
fovi.sensing.projection.angular_directions(cartesian: Tensor, fov_deg: float) → Tensor[source]

Map normalized retinal (N, 2), X right/Y up, to camera unit rays.

fovi.sensing.projection.gaze_rotation(directions: Tensor, convention: str = 'camera_xyz') → Tensor[source]

Aim +Z at (B, 3) directions using the declared zero-torsion convention.

camera_xyz matches Rx(roll) Ry(pitch) in a Y-down/Z-forward camera. pan_tilt matches pan-then-tilt, Ry(pan) Rx(tilt).