{"id":97,"date":"2026-05-08T03:37:45","date_gmt":"2026-05-08T03:37:45","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/?page_id=97"},"modified":"2026-08-26T19:10:29","modified_gmt":"2026-08-26T19:10:29","slug":"related-work","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/","title":{"rendered":"Related Work"},"content":{"rendered":"\n<h2 class=\"wp-block-heading\" id=\"related-work\">3D Visuomotor Imitation Learning<\/h2>\n\n\n\n<p>Recent visuomotor imitation learning methods show that 3D representations are useful for manipulation because point clouds directly encode robot-object geometry. <em><strong><code><a href=\"https:\/\/3d-diffusion-policy.github.io\/\">3D Diffusion Policy (DP3)<\/a><\/code><\/strong><\/em> combines point cloud observations with a diffusion-based action prediction model, making it a strong low-level execution backbone for tasks that require precise geometric control.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"425\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3-1024x425.png\" alt=\"\" class=\"wp-image-141\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3-1024x425.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3-300x125.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3-768x319.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3-1536x637.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3.png 2000w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n\n\n\n<p><strong>In this project, we use a DP3-style low-level policy, but we add explicit goal conditioning. <\/strong>The low-level policy receives the current scene, the observed hand configuration, the predicted goal hand configuration, and the robot state. It then predicts robot actions through a diffusion denoising process.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Goal-Conditioned Hierarchical Policies<\/h3>\n\n\n\n<p><strong><em><code><a href=\"https:\/\/articubot.github.io\/\">ArticuBot<\/a><\/code><\/em><\/strong> demonstrates a hierarchical, goal-conditioned policy for articulated object manipulation with parallel-jaw grippers. Its high-level policy predicts end-effector goal points from the current observation, and its low-level policy executes actions conditioned on the predicted goal. This architecture is effective because it separates interaction-goal prediction from action execution.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1015\" height=\"406\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/image-3.png\" alt=\"\" class=\"wp-image-132\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/image-3.png 1015w, https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/image-3-300x120.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/image-3-768x307.png 768w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n\n\n\n<p><strong>Our project follows this high-level \/ low-level structure, but targets dexterous hands rather than parallel-jaw grippers.<\/strong> This makes the goal design more challenging. A parallel-jaw gripper can often be represented by a few gripper points, while a dexterous hand needs to represent multiple fingers, palm placement, and richer contact geometry.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Dexterous Grasp Priors<\/h3>\n\n\n\n<p>Large-scale dexterous grasp datasets provide useful priors for understanding hand-object interaction. <em><code><strong><a href=\"https:\/\/pku-epic.github.io\/Dexonomy\/\">Dexonomy<\/a><\/strong> <\/code><\/em>provides a taxonomy of dexterous grasp types, helping organize the diversity of hand configurations. <strong><em><code><a href=\"https:\/\/eth-ait.github.io\/graspxl\/\">GraspXL<\/a><\/code><\/em><\/strong> provides large-scale grasping motions for diverse objects, offering another source of dexterous hand-object interaction data.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"473\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/dexonomy_smaller-1024x473.png\" alt=\"\" class=\"wp-image-137\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/dexonomy_smaller-1024x473.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/dexonomy_smaller-300x139.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/dexonomy_smaller-768x355.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/dexonomy_smaller.png 1400w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n\n\n\n<p><strong>These datasets motivate our focus on sparse but expressive hand representations.<\/strong> Rather than representing a goal only as an object pose or a single robot end-effector position, we represent the desired dexterous hand configuration using 3D points on fingers, finger links, and the palm.<\/p>\n\n\n\n<p><\/p>\n\n\n\n<h5 class=\"wp-block-heading\" id=\"related-work\">References<\/h5>\n\n\n<div style=\"font-size: 0.95em;line-height: 1.6\">\n<p>[1] Ze, Yanjie, et al. \u201c3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.\u201d <em>Proceedings of Robotics: Science and Systems<\/em>, 2024. https:\/\/3d-diffusion-policy.github.io\/<\/p>\n<p>[2] Wang, Yufei, et al. \u201cArticuBot: Learning Universal Articulated Object Manipulation Policy via Large Scale Simulation.\u201d <em>Proceedings of Robotics: Science and Systems<\/em>, 2025. https:\/\/articubot.github.io\/<\/p>\n<p>[3] Chen, Jiayi, et al. \u201cDexonomy: Synthesizing All Dexterous Grasp Types in a Grasp Taxonomy.\u201d <em>Proceedings of Robotics: Science and Systems<\/em>, 2025. https:\/\/pku-epic.github.io\/Dexonomy\/<\/p>\n<p>[4] Zhang, Hui, et al. \u201cGraspXL: Generating Grasping Motions for Diverse Objects at Scale.\u201d <em>European Conference on Computer Vision<\/em>, 2024. https:\/\/eth-ait.github.io\/graspxl\/<\/p>\n<\/div>","protected":false},"excerpt":{"rendered":"<p>3D Visuomotor Imitation Learning Recent visuomotor imitation learning methods show that 3D representations are useful for manipulation because point clouds directly encode robot-object geometry. 3D Diffusion Policy (DP3) combines point cloud observations with a diffusion-based action prediction model, making it a strong low-level execution backbone for tasks that require precise geometric control. In this project, &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Related Work&#8221;<\/span><\/a><\/p>\n","protected":false},"author":306,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-97","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Related Work - Goal Conditioning for 3D Dexterous Manipulation Tasks<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Related Work - Goal Conditioning for 3D Dexterous Manipulation Tasks\" \/>\n<meta property=\"og:description\" content=\"3D Visuomotor Imitation Learning Recent visuomotor imitation learning methods show that 3D representations are useful for manipulation because point clouds directly encode robot-object geometry. 3D Diffusion Policy (DP3) combines point cloud observations with a diffusion-based action prediction model, making it a strong low-level execution backbone for tasks that require precise geometric control. In this project, &hellip; Continue reading &quot;Related Work&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/\" \/>\n<meta property=\"og:site_name\" content=\"Goal Conditioning for 3D Dexterous Manipulation Tasks\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-26T19:10:29+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3.png\" \/>\n\t<meta property=\"og:image:width\" content=\"2000\" \/>\n\t<meta property=\"og:image:height\" content=\"830\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"3 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/related-work\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/related-work\\\/\",\"name\":\"Related Work - Goal Conditioning for 3D Dexterous Manipulation Tasks\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/related-work\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/related-work\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/wp-content\\\/uploads\\\/sites\\\/160\\\/2026\\\/05\\\/DP3-1024x425.png\",\"datePublished\":\"2026-05-08T03:37:45+00:00\",\"dateModified\":\"2026-08-26T19:10:29+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/related-work\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/related-work\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/related-work\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/wp-content\\\/uploads\\\/sites\\\/160\\\/2026\\\/05\\\/DP3.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/wp-content\\\/uploads\\\/sites\\\/160\\\/2026\\\/05\\\/DP3.png\",\"width\":2000,\"height\":830},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/related-work\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Related Work\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/\",\"name\":\"Goal Conditioning for 3D Dexterous Manipulation Tasks\",\"description\":\"Xinyu Liu, Advisor: David Held | Robotics Institute, Carnegie Mellon University\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2026teamf18\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Related Work - Goal Conditioning for 3D Dexterous Manipulation Tasks","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/","og_locale":"en_US","og_type":"article","og_title":"Related Work - Goal Conditioning for 3D Dexterous Manipulation Tasks","og_description":"3D Visuomotor Imitation Learning Recent visuomotor imitation learning methods show that 3D representations are useful for manipulation because point clouds directly encode robot-object geometry. 3D Diffusion Policy (DP3) combines point cloud observations with a diffusion-based action prediction model, making it a strong low-level execution backbone for tasks that require precise geometric control. In this project, &hellip; Continue reading \"Related Work\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/","og_site_name":"Goal Conditioning for 3D Dexterous Manipulation Tasks","article_modified_time":"2026-08-26T19:10:29+00:00","og_image":[{"width":2000,"height":830,"url":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3.png","type":"image\/png"}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"3 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/","name":"Related Work - Goal Conditioning for 3D Dexterous Manipulation Tasks","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3-1024x425.png","datePublished":"2026-05-08T03:37:45+00:00","dateModified":"2026-08-26T19:10:29+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-content\/uploads\/sites\/160\/2026\/05\/DP3.png","width":2000,"height":830},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/related-work\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/"},{"@type":"ListItem","position":2,"name":"Related Work"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/","name":"Goal Conditioning for 3D Dexterous Manipulation Tasks","description":"Xinyu Liu, Advisor: David Held | Robotics Institute, Carnegie Mellon University","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-json\/wp\/v2\/pages\/97","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-json\/wp\/v2\/users\/306"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-json\/wp\/v2\/comments?post=97"}],"version-history":[{"count":16,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-json\/wp\/v2\/pages\/97\/revisions"}],"predecessor-version":[{"id":191,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-json\/wp\/v2\/pages\/97\/revisions\/191"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2026teamf18\/wp-json\/wp\/v2\/media?parent=97"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}