{"id":46,"date":"2023-05-10T00:00:59","date_gmt":"2023-05-10T00:00:59","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/?page_id=46"},"modified":"2023-12-18T06:43:21","modified_gmt":"2023-12-18T06:43:21","slug":"methodology","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/","title":{"rendered":"Methodology"},"content":{"rendered":"\n<p><strong>TLDR:<\/strong> We model the HOI scene by a time-persistent implicit field for the object, hand meshes parameterized by hand shape, hand articulation, along with a time-varying rigid transformation for object pose. We define the cameras in the hand frame. We optimize a video-specific scene representation using reprojection loss from the original view and diffusion distillation loss from a novel view. Our diffusion model takes in a noisy geometry rendering of the object, the geometry rendering of the hand, and a text prompt, to output the denoised geometry rendering of objects.<\/p>\n\n\n\n<p><strong>Overview:<\/strong><\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"240\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-1024x240.png\" alt=\"\" class=\"wp-image-47\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-1024x240.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-300x70.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-768x180.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-1536x359.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-2048x479.png 2048w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n\n\n\n<p class=\"has-text-align-center\"><strong>Method Figure:<\/strong> We compose the scene at time step t such that it can be reprojected back to the image space M from the camera pose v<sub>t<\/sub>. With an arbitrary view point v\u2032, we then render amodal geometry cues for the object (surface normal N<sub>o<\/sub>, depth D<sub>o<\/sub>, and mask M<sub>o<\/sub>) as well as the hand. We expect the hand rendering H and an optional text condition on the object category, C, to be predictive of the object rendering O.<\/p>\n\n\n\n<p>We first represent the HOI scene as hand and the object seperately. We have a time persistent Implicit field for the Rigid Object that can handle unknown object topologies. DeepSDF MLP layers are used to predict the Signed Distance Function (SDF to the object surface <\/p>\n\n\n\n<p class=\"has-text-align-center\">\u03d5(X) = s. <\/p>\n\n\n\n<p>Parallely, we also have time varying Hand meshes at every time step t of the video. Given the texture of the hand (\u03b2) and its articulation (\u03b8<sup>t<\/sup><sub>A<\/sub>) which are 10-dimensional and 45-dimensional respectively, we use a predefined parametric mesh model (MANO) to represent the hand dynamics.<\/p>\n\n\n\n<p class=\"has-text-align-center\">H<sup>t<\/sup> = MANO(\u03b2, \u03b8<sup>t<\/sup><sub>A<\/sub>)<\/p>\n\n\n\n<p>Given the time-persistent object representation \u03d5 and a time-varying hand mesh H<sup>t<\/sup>, we then compose them into a scene at time t such that they can be reprojected back to the image space from the cameras. Since we don\u2019t have access to the object templates, we track object pose with respect to hand T<sup>t<\/sup><sub>h\u2192o<\/sub> and initialize them to identity for the first frame. The rigid object frame can be related to the predicted camera frame by composing the two transformations: T<sup>t<\/sup><sub>c\u2192h<\/sub> and T<sup>t<\/sup><sub>h\u2192o<\/sub>.<\/p>\n\n\n\n<p>To differentiably render the HOI scene, we seperately render the object using volumentric rendering and the hand using mesh rendering to obtain 3 geometric cues: Normal, Depth and Mask, i.e G<sub>o<\/sub> \u2261 (M<sub>o<\/sub>,D<sub>o<\/sub>,N<sub>o<\/sub>), G<sub>h<\/sub> \u2261 (M<sub>h<\/sub>,D<sub>h<\/sub>,N<sub>h<\/sub>). To compose them into semantic masks, we then blend these renderings into HOI images by their predicted rendered depth: <\/p>\n\n\n\n<p class=\"has-text-align-center\">M = B(M<sub>h<\/sub>,M<sub>o<\/sub>,D<sub>h<\/sub>,D<sub>o<\/sub>)<\/p>\n\n\n\n<p><strong>Loss functions:<\/strong><\/p>\n\n\n\n<p><em>1. Hoi representation needs to be optimized to explain the input sequence. We therefore render semantic mask<\/em> <em>from estimated cameras (every frame) and compare the reprojection error from the the estimated original views<\/em> <em>with ground truth masks: <\/em><\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"177\" height=\"20\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-14.png\" alt=\"\" class=\"wp-image-49\" \/><\/figure>\n<\/div>\n\n\n<p>2. We need to optimize the scene to appear more likely from a novel view point. Scored distillation sampling treats the output of diffusion model as a critic to approximate towards more likely images (without backpropogating for efficiency). We therefore apply a loss on the reconstructed denoised signal G<sup>i<\/sup><sub>o<\/sub> from the pretrained diffusion <em>model: <\/em><\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-15.png\" alt=\"\" class=\"wp-image-50\" width=\"241\" height=\"33\" \/><\/figure>\n<\/div>\n\n\n<p>where v is a novel camera view, \u03f5 is the noise, and i is the time step.<\/p>\n\n\n\n<p><strong>Data Driven Prior for HOI Geometry<\/strong>:<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"635\" height=\"234\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-16.png\" alt=\"\" class=\"wp-image-51\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-16.png 635w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-16-300x111.png 300w\" sizes=\"auto, (max-width: 635px) 100vw, 635px\" \/><\/figure>\n\n\n\n<p class=\"has-text-align-center\"><strong>Figure:<\/strong> <strong>Geometry-informed Diffusion<\/strong>: Our diffusion model aims to take in a noisy geometry rendering of the object, the geometry rendering of the hand, and a text prompt, to output the denoised geometry rendering of objects.<\/p>\n\n\n\n<p>We want to condition the diffusion model on category cues of an object (such as &#8220;cylindrical mug&#8221;) and hand cues (such as pinched hands imply thin handles). The diffusion model learns a data-driven distribution over geometry rendering of objects given its category C and the interacting hand H since it\u2019s pretrained with large-scale ground truth HOIs. We use this learned prior to guiding per-sequence optimization and we expect the diffusion model to capture the likelihood of a common object geometry given C, H: p(\u03d5<sup>t<\/sup>|H,C). The diffusion prior basically helps to ensure that the novel views of the object are reasonable.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>TLDR: We model the HOI scene by a time-persistent implicit field for the object, hand meshes parameterized by hand shape, hand articulation, along with a time-varying rigid transformation for object pose. We define the cameras in the hand frame. We optimize a video-specific scene representation using reprojection loss from the original view and diffusion distillation &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Methodology&#8221;<\/span><\/a><\/p>\n","protected":false},"author":186,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-46","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Methodology - Reconstructing Hand-Object Interactions from Internet Videos<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Methodology - Reconstructing Hand-Object Interactions from Internet Videos\" \/>\n<meta property=\"og:description\" content=\"TLDR: We model the HOI scene by a time-persistent implicit field for the object, hand meshes parameterized by hand shape, hand articulation, along with a time-varying rigid transformation for object pose. We define the cameras in the hand frame. We optimize a video-specific scene representation using reprojection loss from the original view and diffusion distillation &hellip; Continue reading &quot;Methodology&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/\" \/>\n<meta property=\"og:site_name\" content=\"Reconstructing Hand-Object Interactions from Internet Videos\" \/>\n<meta property=\"article:modified_time\" content=\"2023-12-18T06:43:21+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-1024x240.png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"4 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/methodology\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/methodology\\\/\",\"name\":\"Methodology - Reconstructing Hand-Object Interactions from Internet Videos\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/methodology\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/methodology\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/wp-content\\\/uploads\\\/sites\\\/95\\\/2023\\\/05\\\/image-12-1024x240.png\",\"datePublished\":\"2023-05-10T00:00:59+00:00\",\"dateModified\":\"2023-12-18T06:43:21+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/methodology\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/methodology\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/methodology\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/wp-content\\\/uploads\\\/sites\\\/95\\\/2023\\\/05\\\/image-12-1024x240.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/wp-content\\\/uploads\\\/sites\\\/95\\\/2023\\\/05\\\/image-12-1024x240.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/methodology\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Methodology\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/\",\"name\":\"Reconstructing Hand-Object Interactions from Internet Videos\",\"description\":\"\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Methodology - Reconstructing Hand-Object Interactions from Internet Videos","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/","og_locale":"en_US","og_type":"article","og_title":"Methodology - Reconstructing Hand-Object Interactions from Internet Videos","og_description":"TLDR: We model the HOI scene by a time-persistent implicit field for the object, hand meshes parameterized by hand shape, hand articulation, along with a time-varying rigid transformation for object pose. We define the cameras in the hand frame. We optimize a video-specific scene representation using reprojection loss from the original view and diffusion distillation &hellip; Continue reading \"Methodology\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/","og_site_name":"Reconstructing Hand-Object Interactions from Internet Videos","article_modified_time":"2023-12-18T06:43:21+00:00","og_image":[{"url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-1024x240.png","type":"","width":"","height":""}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/","name":"Methodology - Reconstructing Hand-Object Interactions from Internet Videos","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-1024x240.png","datePublished":"2023-05-10T00:00:59+00:00","dateModified":"2023-12-18T06:43:21+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-1024x240.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-12-1024x240.png"},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/methodology\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/"},{"@type":"ListItem","position":2,"name":"Methodology"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/","name":"Reconstructing Hand-Object Interactions from Internet Videos","description":"","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages\/46","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/users\/186"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/comments?post=46"}],"version-history":[{"count":7,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages\/46\/revisions"}],"predecessor-version":[{"id":197,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages\/46\/revisions\/197"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/media?parent=46"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}