Hi, thanks for the great work and for releasing EgoInfinity.
I have a question about the synthesized egocentric camera placement described in Appendix A.1.
From the paper, my understanding is that for each frame, the virtual ego camera is placed at a fixed offset above a hand-derived anchor, specifically the bilateral hand midpoint, with its up-axis aligned to the GeoCalib gravity vector, and then aimed at the hand-object interaction region.
However, in the demo/visualization, the synthesized egocentric view looks like the camera may be offset in two directions, e.g. both upward along gravity and backward relative to the interaction/workspace direction. If the camera were offset only along the gravity/up direction, I would expect the view to look more top-down.
Could you clarify whether the synthesized ego camera uses only a gravity/up-direction offset, or whether there is also an additional backward/forward offset direction? If the latter, how is that direction computed, and what offset values are used?
I may have missed this detail in the paper, so I would appreciate any clarification.
Hi, thanks for the great work and for releasing EgoInfinity.
I have a question about the synthesized egocentric camera placement described in Appendix A.1.
From the paper, my understanding is that for each frame, the virtual ego camera is placed at a fixed offset above a hand-derived anchor, specifically the bilateral hand midpoint, with its up-axis aligned to the GeoCalib gravity vector, and then aimed at the hand-object interaction region.
However, in the demo/visualization, the synthesized egocentric view looks like the camera may be offset in two directions, e.g. both upward along gravity and backward relative to the interaction/workspace direction. If the camera were offset only along the gravity/up direction, I would expect the view to look more top-down.
Could you clarify whether the synthesized ego camera uses only a gravity/up-direction offset, or whether there is also an additional backward/forward offset direction? If the latter, how is that direction computed, and what offset values are used?
I may have missed this detail in the paper, so I would appreciate any clarification.