{"id":2,"date":"2023-05-01T17:49:07","date_gmt":"2023-05-01T17:49:07","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/?page_id=2"},"modified":"2023-12-18T06:43:17","modified_gmt":"2023-12-18T06:43:17","slug":"introduction","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/","title":{"rendered":"Introduction"},"content":{"rendered":"\n<p>We all, in general, not only run and work, but also interact with the object around us. Like playing tennis, drinking tea etc. Instead of just reconstructing the human and the hand, we also want to reconstruct the racket and tennis ball that the player holds, the mouse and pen in someone\u2019s hands. In particular, we need to understand the object in hand and the type of interaction between the hand and the object. In this project, we focus on hand-object reconstruction. This is a very challenging problem because hands and objects often suffer from heavy mutual-occlusion, so reconstructing them could be quite difficult.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-10.png\" alt=\"\" class=\"wp-image-40\" width=\"433\" height=\"348\" \/><\/figure>\n<\/div>\n\n\n<p>Imagine watching a video of human cooking in the kitchen. We humans can easily understand the underlying 3D hand-object-interactions in the physical world, such as contact regions, 3D relations between hands and objects. More impressively, we are able to hallucinate the 3D shape even if we have not seen some views of the object, such as the bottom part of a cooking pan. <\/p>\n\n\n\n<p><\/p>\n\n\n\n<p>In this work, we aim to reconstruct 3D hand-object interactions (HOI) from a monocular video. Prior works of HOI from videos typically assume known object templates and jointly fit hand and object pose to the 3D instance. In contrast, we aim to reconstruct unknown generic objects. To do so, on one hand, many works have studied 3D reconstruction for generic scenes by optimizing a neural field to match the given image collections or video. However, they typically require all object regions visible in some of the frames. On the other hand, recent works have studied a data-driven approach to learn hand-object prior but they do not consider temporal consistency in videos. In this work, we propose to leverage progress from both directions to reconstruct the hand and object from Internet video clips. We first explore reconstruction for hand-held rigid object and then try to extend the architecture to articulating objects.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>We all, in general, not only run and work, but also interact with the object around us. Like playing tennis, drinking tea etc. Instead of just reconstructing the human and the hand, we also want to reconstruct the racket and tennis ball that the player holds, the mouse and pen in someone\u2019s hands. In particular, &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Introduction&#8221;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-2","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Introduction - Reconstructing Hand-Object Interactions from Internet Videos<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Introduction - Reconstructing Hand-Object Interactions from Internet Videos\" \/>\n<meta property=\"og:description\" content=\"We all, in general, not only run and work, but also interact with the object around us. Like playing tennis, drinking tea etc. Instead of just reconstructing the human and the hand, we also want to reconstruct the racket and tennis ball that the player holds, the mouse and pen in someone\u2019s hands. In particular, &hellip; Continue reading &quot;Introduction&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/\" \/>\n<meta property=\"og:site_name\" content=\"Reconstructing Hand-Object Interactions from Internet Videos\" \/>\n<meta property=\"article:modified_time\" content=\"2023-12-18T06:43:17+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-10.png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"2 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/\",\"name\":\"Introduction - Reconstructing Hand-Object Interactions from Internet Videos\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/wp-content\\\/uploads\\\/sites\\\/95\\\/2023\\\/05\\\/image-10.png\",\"datePublished\":\"2023-05-01T17:49:07+00:00\",\"dateModified\":\"2023-12-18T06:43:17+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/wp-content\\\/uploads\\\/sites\\\/95\\\/2023\\\/05\\\/image-10.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/wp-content\\\/uploads\\\/sites\\\/95\\\/2023\\\/05\\\/image-10.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Introduction\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/\",\"name\":\"Reconstructing Hand-Object Interactions from Internet Videos\",\"description\":\"\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team18\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Introduction - Reconstructing Hand-Object Interactions from Internet Videos","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/","og_locale":"en_US","og_type":"article","og_title":"Introduction - Reconstructing Hand-Object Interactions from Internet Videos","og_description":"We all, in general, not only run and work, but also interact with the object around us. Like playing tennis, drinking tea etc. Instead of just reconstructing the human and the hand, we also want to reconstruct the racket and tennis ball that the player holds, the mouse and pen in someone\u2019s hands. In particular, &hellip; Continue reading \"Introduction\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/","og_site_name":"Reconstructing Hand-Object Interactions from Internet Videos","article_modified_time":"2023-12-18T06:43:17+00:00","og_image":[{"url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-10.png","type":"","width":"","height":""}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"2 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/","name":"Introduction - Reconstructing Hand-Object Interactions from Internet Videos","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-10.png","datePublished":"2023-05-01T17:49:07+00:00","dateModified":"2023-12-18T06:43:17+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-10.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-content\/uploads\/sites\/95\/2023\/05\/image-10.png"},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/"},{"@type":"ListItem","position":2,"name":"Introduction"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/","name":"Reconstructing Hand-Object Interactions from Internet Videos","description":"","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages\/2","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/comments?post=2"}],"version-history":[{"count":5,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages\/2\/revisions"}],"predecessor-version":[{"id":149,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/pages\/2\/revisions\/149"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team18\/wp-json\/wp\/v2\/media?parent=2"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}