{"id":15,"date":"2023-05-09T22:18:23","date_gmt":"2023-05-09T22:18:23","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/?page_id=15"},"modified":"2023-12-18T06:43:36","modified_gmt":"2023-12-18T06:43:36","slug":"task-formulation","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/","title":{"rendered":"Task Formulation"},"content":{"rendered":"\n<div class=\"wp-block-columns is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex\">\n<div class=\"wp-block-column is-layout-flow wp-block-column-is-layout-flow\" style=\"flex-basis:33.33%\">\n<figure class=\"wp-block-video\"><video height=\"512\" style=\"aspect-ratio: 512 \/ 512;\" width=\"512\" controls src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/tf3.mp4\"><\/video><\/figure>\n<\/div>\n\n\n\n<div class=\"wp-block-column is-layout-flow wp-block-column-is-layout-flow\" style=\"flex-basis:66.66%\">\n<figure class=\"wp-block-video aligncenter\"><video height=\"200\" style=\"aspect-ratio: 800 \/ 200;\" width=\"800\" controls src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/tf4.mp4\"><\/video><\/figure>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-columns is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex\">\n<div class=\"wp-block-column is-layout-flow wp-block-column-is-layout-flow\" style=\"flex-basis:33.33%\">\n<figure class=\"wp-block-video\"><video height=\"512\" style=\"aspect-ratio: 512 \/ 512;\" width=\"512\" controls src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/tf1-2.mp4\"><\/video><\/figure>\n<\/div>\n\n\n\n<div class=\"wp-block-column is-layout-flow wp-block-column-is-layout-flow\" style=\"flex-basis:66.66%\">\n<figure class=\"wp-block-video aligncenter\"><video height=\"200\" style=\"aspect-ratio: 800 \/ 200;\" width=\"800\" controls src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/tf2-1.mp4\"><\/video><\/figure>\n<\/div>\n<\/div>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"381\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png\" alt=\"\" class=\"wp-image-165\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1-300x112.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1-768x285.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1.png 1426w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n\n\n\n<p>Given a video clip depicting a hand interacting with a rigid object, we aim to infer the underlying 3D shape of both the hand and the object, i.e:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>3D shape of the object (\u03d5)<\/li>\n\n\n\n<li>Texture of the hand (\u03b2)<\/li>\n\n\n\n<li>Intrinsics of the camera (K<sup>t<\/sup>)<\/li>\n\n\n\n<li>Per frame articulation of the hand (\u03b8<sup>t<\/sup><sub>A<\/sub>)<\/li>\n\n\n\n<li>Per frame object pose (T<sub>h\u2192o<\/sub>)<\/li>\n\n\n\n<li>Per frame camera pose (T<sub>c\u2192h<\/sub>)<\/li>\n<\/ul>\n\n\n\n<p class=\"has-text-align-center\"><strong>Idea: Hand a visual cue for Object Shape<\/strong><\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-8-1024x363.png\" alt=\"\" class=\"wp-image-35\" width=\"347\" height=\"122\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-8-1024x363.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-8-300x106.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-8-768x272.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-8.png 1088w\" sizes=\"auto, (max-width: 347px) 100vw, 347px\" \/><\/figure>\n<\/div>\n\n\n<p>Hand can actually be thought of as a special occluder. There are 2 reasons behind this idea:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Hand pose is predictive of the object\u2019s shape. For example, When we pinch our fingers together, we are very likely to hold thin sticks, (like pens, nails, and brushes.)<\/li>\n\n\n\n<li>The current hand pose reconstruction system is quite robust to occlusion and their prediction can often be trusted.<\/li>\n<\/ul>\n\n\n\n<p>Given an image (a single frame), we therefore try to simultaneously optimize 2 things for a good reconstruction: i) Estimate the underlying hand pose and ii) infer the object shape in a normalized hand-centric coordinate frame. We anticipate explicit articulation conditioning to help the inference. Extending it to a dynamic setting, we intend to consider the Neural Field of the scene and fit a prior about the understanding of the possible valid shapes of the object.<\/p>\n\n\n\n<p>Given the model, we can then get a per-frame prediction of the interaction and since the NeRF belongs to a valid human-object pair, we expect the frame predictions to be consistent.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Given a video clip depicting a hand interacting with a rigid object, we aim to infer the underlying 3D shape of both the hand and the object, i.e: Idea: Hand a visual cue for Object Shape Hand can actually be thought of as a special occluder. There are 2 reasons behind this idea: Given an &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Task Formulation&#8221;<\/span><\/a><\/p>\n","protected":false},"author":186,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-15","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Task Formulation - Reconstructing Hand-Object Interactions from Internet Videos<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Task Formulation - Reconstructing Hand-Object Interactions from Internet Videos\" \/>\n<meta property=\"og:description\" content=\"Given a video clip depicting a hand interacting with a rigid object, we aim to infer the underlying 3D shape of both the hand and the object, i.e: Idea: Hand a visual cue for Object Shape Hand can actually be thought of as a special occluder. There are 2 reasons behind this idea: Given an &hellip; Continue reading &quot;Task Formulation&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/\" \/>\n<meta property=\"og:site_name\" content=\"Reconstructing Hand-Object Interactions from Internet Videos\" \/>\n<meta property=\"article:modified_time\" content=\"2023-12-18T06:43:36+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"2 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/task-formulation\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/task-formulation\\\/\",\"name\":\"Task Formulation - Reconstructing Hand-Object Interactions from Internet Videos\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/task-formulation\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/task-formulation\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/wp-content\\\/uploads\\\/sites\\\/95\\\/2023\\\/05\\\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png\",\"datePublished\":\"2023-05-09T22:18:23+00:00\",\"dateModified\":\"2023-12-18T06:43:36+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/task-formulation\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/task-formulation\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/task-formulation\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/wp-content\\\/uploads\\\/sites\\\/95\\\/2023\\\/05\\\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/wp-content\\\/uploads\\\/sites\\\/95\\\/2023\\\/05\\\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/task-formulation\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Task Formulation\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/\",\"name\":\"Reconstructing Hand-Object Interactions from Internet Videos\",\"description\":\"\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Task Formulation - Reconstructing Hand-Object Interactions from Internet Videos","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/","og_locale":"en_US","og_type":"article","og_title":"Task Formulation - Reconstructing Hand-Object Interactions from Internet Videos","og_description":"Given a video clip depicting a hand interacting with a rigid object, we aim to infer the underlying 3D shape of both the hand and the object, i.e: Idea: Hand a visual cue for Object Shape Hand can actually be thought of as a special occluder. There are 2 reasons behind this idea: Given an &hellip; Continue reading \"Task Formulation\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/","og_site_name":"Reconstructing Hand-Object Interactions from Internet Videos","article_modified_time":"2023-12-18T06:43:36+00:00","og_image":[{"url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png","type":"","width":"","height":""}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"2 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/","name":"Task Formulation - Reconstructing Hand-Object Interactions from Internet Videos","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png","datePublished":"2023-05-09T22:18:23+00:00","dateModified":"2023-12-18T06:43:36+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/Screenshot-2023-05-09-at-10.10.16-PM-1-1024x381.png"},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/task-formulation\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/"},{"@type":"ListItem","position":2,"name":"Task Formulation"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/","name":"Reconstructing Hand-Object Interactions from Internet Videos","description":"","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages\/15","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/users\/186"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/comments?post=15"}],"version-history":[{"count":7,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages\/15\/revisions"}],"predecessor-version":[{"id":200,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages\/15\/revisions\/200"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/media?parent=15"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}