{"id":7,"date":"2025-05-07T18:31:51","date_gmt":"2025-05-07T18:31:51","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/?page_id=7"},"modified":"2025-12-11T05:12:09","modified_gmt":"2025-12-11T05:12:09","slug":"methodology","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/","title":{"rendered":"Methodologies"},"content":{"rendered":"\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-constrained wp-block-group-is-layout-constrained\">\n<div class=\"wp-block-columns is-layout-flex wp-container-core-columns-is-layout-9d6595d7 wp-block-columns-is-layout-flex\">\n<div class=\"wp-block-column is-layout-flow wp-block-column-is-layout-flow\" style=\"flex-basis:100%\"><\/div>\n<\/div>\n<\/div><\/div>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"635\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-1024x635.png\" alt=\"\" class=\"wp-image-156\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-1024x635.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-300x186.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-768x476.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-1536x953.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-2048x1270.png 2048w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\">Fig. 1 &#8211; Two Pipelines<\/figcaption><\/figure>\n\n\n\n<p>Upon discussions with our project sponsor, we converged on two main pipelines that would allow the user to easier generate pipelines. <\/p>\n\n\n\n<p>The first, shown on top, is <strong>Control via Demonstration<\/strong> &#8211; a method by which the user can generate a trajectory by recording it on their phone, and the robotic camera rig will mimic this output to the best of its ability, while recognizing its physical constraints. Further details are included in the corresponding page.<\/p>\n\n\n\n<p>The second, shown on the bottom, is <strong>Control via Natural Cinematographic Prompts<\/strong> &#8211; a method where the user can either type or speak a prompt directing what the shot will look like, and the system will generate a 3D grounded trajectory that the robotic rig can then follow. This is done by our Intent-to-Motion Layer, which builds upon the methodology introduced by Liu et al. in 2024 [1]. It has three key components: LLM Agent, a Text-To-Trajectory model, and an anchor detector. Further details are included in the corresponding page.<\/p>\n\n\n\n<p>A final <strong>Trajectory Refinement Layer<\/strong> is used to ensure the trajectory fits within physical constrains of the robotic rig, while maintaining the directorial intent behind the input or generated trajectory.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Trajectory Refinement and Execution Layer<\/h2>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"alignright size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"908\" height=\"1024\" src=\"http:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/Trajectory-Refinement-908x1024.png\" alt=\"\" class=\"wp-image-47\" style=\"width:356px;height:auto\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/Trajectory-Refinement-908x1024.png 908w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/Trajectory-Refinement-266x300.png 266w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/Trajectory-Refinement-768x866.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/Trajectory-Refinement-1362x1536.png 1362w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/Trajectory-Refinement-1815x2048.png 1815w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\">Fig. 2 &#8211; Overview of modules in Trajectory Refinement and Execution Layer<\/figcaption><\/figure>\n<\/div>\n\n\n<p>Trajectory generation may result in a non-smooth trajectory with rapid angle and speed changes. This module is responsible for generating a smooth trajectory based on predefined heuristics.<\/p>\n\n\n\n<p>Once we have a trajectory, we need to convert it to control commands that can be understood by the robotic rig. <br><br>The inverse kinematics module solves for various control values given the 3D position in the world coordinate system. <br><br>We assume a 7-DOF cinematic robot arm as shown in Figure 3.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"alignleft size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"788\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/robot-controls-1-1024x788.jpg\" alt=\"\" class=\"wp-image-77\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/robot-controls-1-1024x788.jpg 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/robot-controls-1-300x231.jpg 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/robot-controls-1-768x591.jpg 768w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/robot-controls-1-1536x1182.jpg 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/robot-controls-1.jpg 1686w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\">Fig. 3 &#8211; Control Names for Various Parts of the Camera Rig<\/figcaption><\/figure>\n<\/div>\n\n\n<p>Assuming dynamics of a 7-DOF cinematic robotic arm, we solve for the values for 4 control joints (Arm, Lift, Rotate, and Track) by solving the set of equations below in a least squares manner. While many of these are accounted for by Flair, we have to solve for the Arm value ourselves, necessitating this inverse kinematics model.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"292\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/image-3-1024x292.png\" alt=\"\" class=\"wp-image-88\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/image-3-1024x292.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/image-3-300x85.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/image-3-768x219.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/05\/image-3.png 1334w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">References<\/h2>\n\n\n\n<p>[1] Xinhang Liu, Yu-Wing Tai, and Chi-Keung Tang. ChatCam: Empowering camera control through conversational AI, 2024.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Upon discussions with our project sponsor, we converged on two main pipelines that would allow the user to easier generate pipelines. The first, shown on top, is Control via Demonstration &#8211; a method by which the user can generate a trajectory by recording it on their phone, and the robotic camera rig will mimic this &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Methodologies&#8221;<\/span><\/a><\/p>\n","protected":false},"author":251,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-7","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Methodologies - Computer Vision for Cinematographic Motion Control<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Methodologies - Computer Vision for Cinematographic Motion Control\" \/>\n<meta property=\"og:description\" content=\"Upon discussions with our project sponsor, we converged on two main pipelines that would allow the user to easier generate pipelines. The first, shown on top, is Control via Demonstration &#8211; a method by which the user can generate a trajectory by recording it on their phone, and the robotic camera rig will mimic this &hellip; Continue reading &quot;Methodologies&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/\" \/>\n<meta property=\"og:site_name\" content=\"Computer Vision for Cinematographic Motion Control\" \/>\n<meta property=\"article:modified_time\" content=\"2025-12-11T05:12:09+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-scaled.png\" \/>\n\t<meta property=\"og:image:width\" content=\"2560\" \/>\n\t<meta property=\"og:image:height\" content=\"1588\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"3 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/methodology\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/methodology\\\/\",\"name\":\"Methodologies - Computer Vision for Cinematographic Motion Control\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/methodology\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/methodology\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/wp-content\\\/uploads\\\/sites\\\/133\\\/2025\\\/12\\\/CC_archi_fixed-1024x635.png\",\"datePublished\":\"2025-05-07T18:31:51+00:00\",\"dateModified\":\"2025-12-11T05:12:09+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/methodology\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/methodology\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/methodology\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/wp-content\\\/uploads\\\/sites\\\/133\\\/2025\\\/12\\\/CC_archi_fixed-scaled.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/wp-content\\\/uploads\\\/sites\\\/133\\\/2025\\\/12\\\/CC_archi_fixed-scaled.png\",\"width\":2560,\"height\":1588},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/methodology\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Methodologies\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/\",\"name\":\"Computer Vision for Cinematographic Motion Control\",\"description\":\"Shaurye Aggarwal and Kaustav Mukherjee\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team1\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Methodologies - Computer Vision for Cinematographic Motion Control","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/","og_locale":"en_US","og_type":"article","og_title":"Methodologies - Computer Vision for Cinematographic Motion Control","og_description":"Upon discussions with our project sponsor, we converged on two main pipelines that would allow the user to easier generate pipelines. The first, shown on top, is Control via Demonstration &#8211; a method by which the user can generate a trajectory by recording it on their phone, and the robotic camera rig will mimic this &hellip; Continue reading \"Methodologies\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/","og_site_name":"Computer Vision for Cinematographic Motion Control","article_modified_time":"2025-12-11T05:12:09+00:00","og_image":[{"width":2560,"height":1588,"url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-scaled.png","type":"image\/png"}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"3 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/","name":"Methodologies - Computer Vision for Cinematographic Motion Control","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-1024x635.png","datePublished":"2025-05-07T18:31:51+00:00","dateModified":"2025-12-11T05:12:09+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-scaled.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-content\/uploads\/sites\/133\/2025\/12\/CC_archi_fixed-scaled.png","width":2560,"height":1588},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/methodology\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/"},{"@type":"ListItem","position":2,"name":"Methodologies"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/","name":"Computer Vision for Cinematographic Motion Control","description":"Shaurye Aggarwal and Kaustav Mukherjee","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-json\/wp\/v2\/pages\/7","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-json\/wp\/v2\/users\/251"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-json\/wp\/v2\/comments?post=7"}],"version-history":[{"count":29,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-json\/wp\/v2\/pages\/7\/revisions"}],"predecessor-version":[{"id":157,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-json\/wp\/v2\/pages\/7\/revisions\/157"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team1\/wp-json\/wp\/v2\/media?parent=7"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}