{"id":71,"date":"2025-12-10T05:06:39","date_gmt":"2025-12-10T05:06:39","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/?page_id=71"},"modified":"2025-12-10T05:25:49","modified_gmt":"2025-12-10T05:25:49","slug":"experiment-results","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/","title":{"rendered":"Experiment Results"},"content":{"rendered":"\n<p><strong>Comparisons<\/strong><br>As shown, our method achieves approximately 3\u00d7 faster inference compared to existing audio-driven and pose-driven baselines. In addition to speed, our approach produces higher-quality and more realistic results. Compared with audio-driven methods, our model not only maintains high generation quality but also substantially improves Sync-C and HKC. In particular, the lip synchronization confidence increases from 4.36 to 7.26. Moreover, compared with S2G-MDD, our method improves HKC from 0.956 to 0.968 on the test set.<\/p>\n\n\n\n<p>Compared with pose-driven methods, our approach outperforms all baselines in both lip synchronization and overall motion quality. Remarkably, our student model, when compared to its teacher model MimicMotion, achieves a 13.1\u00d7 inference speed-up without sacrificing generation quality, while also further improving motion and synchronization metrics. Specifically, our method improves HKC from 0.928 to 0.948 and Sync-C from 4.56 to 7.28, demonstrating enhanced hand motion confidence and lip synchronization performance.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"832\" height=\"384\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.18.45-AM.png\" alt=\"\" class=\"wp-image-75\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.18.45-AM.png 832w, https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.18.45-AM-300x138.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.18.45-AM-768x354.png 768w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"820\" height=\"546\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.19.09-AM.png\" alt=\"\" class=\"wp-image-76\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.19.09-AM.png 820w, https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.19.09-AM-300x200.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.19.09-AM-768x511.png 768w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n\n\n\n<p><\/p>\n\n\n\n<p><strong>Ablation Studies<\/strong><\/p>\n\n\n\n<p>We conduct an ablation study to analyze the contribution of each component. Using the teacher model as the baseline, we observe strong motion quality and lip synchronization performance after fine-tuning on the co-speech dataset. However, its inference speed remains a bottleneck at 1.93 FPS.Both input-aware global attention and input-aware local attention preserve generation quality without degradation, but the speed improvements they bring are limited. Directly applying DMD distillation provides a substantial speedup but introduces noticeable quality degradation, especially artifacts on faces and hands.By incorporating our input-aware distillation strategy, the model achieves real-time performance at 25.31 FPS while maintaining generation quality comparable to the fine-tuned teacher model.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"810\" height=\"860\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.23.51-AM.png\" alt=\"\" class=\"wp-image-77\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.23.51-AM.png 810w, https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.23.51-AM-283x300.png 283w, https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.23.51-AM-768x815.png 768w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n\n\n\n<p>We present a detailed breakdown of the inference time across different architectural variants by decomposing the total runtime into four components: Attention, Linear, Norm, and Others. The baseline teacher model requires 103.6 seconds to process an 8-second video. By introducing global attention, we reduce the processing time to 60.9 seconds, primarily due to reductions in attention and linear computation. Adding local attention further decreases the time to 45.2 seconds. Finally, applying distillation reduces the runtime to 7.9 seconds, achieving a 13.1\u00d7 speedup compared to the teacher model, largely driven by the substantial reduction in attention cost. This demonstrates the effectiveness of our sparse attention strategy in enabling real-time co-speech video generation.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"890\" height=\"418\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.24.02-AM.png\" alt=\"\" class=\"wp-image-78\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.24.02-AM.png 890w, https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.24.02-AM-300x141.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.24.02-AM-768x361.png 768w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><\/figure>\n","protected":false},"excerpt":{"rendered":"<p>ComparisonsAs shown, our method achieves approximately 3\u00d7 faster inference compared to existing audio-driven and pose-driven baselines. In addition to speed, our approach produces higher-quality and more realistic results. Compared with audio-driven methods, our model not only maintains high generation quality but also substantially improves Sync-C and HKC. In particular, the lip synchronization confidence increases from &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Experiment Results&#8221;<\/span><\/a><\/p>\n","protected":false},"author":260,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-71","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Experiment Results - 2025 Team 16-1 Project<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Experiment Results - 2025 Team 16-1 Project\" \/>\n<meta property=\"og:description\" content=\"ComparisonsAs shown, our method achieves approximately 3\u00d7 faster inference compared to existing audio-driven and pose-driven baselines. In addition to speed, our approach produces higher-quality and more realistic results. Compared with audio-driven methods, our model not only maintains high generation quality but also substantially improves Sync-C and HKC. In particular, the lip synchronization confidence increases from &hellip; Continue reading &quot;Experiment Results&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/\" \/>\n<meta property=\"og:site_name\" content=\"2025 Team 16-1 Project\" \/>\n<meta property=\"article:modified_time\" content=\"2025-12-10T05:25:49+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.18.45-AM.png\" \/>\n\t<meta property=\"og:image:width\" content=\"832\" \/>\n\t<meta property=\"og:image:height\" content=\"384\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"3 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/experiment-results\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/experiment-results\\\/\",\"name\":\"Experiment Results - 2025 Team 16-1 Project\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/experiment-results\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/experiment-results\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/wp-content\\\/uploads\\\/sites\\\/138\\\/2025\\\/12\\\/Screenshot-2025-12-10-at-12.18.45-AM.png\",\"datePublished\":\"2025-12-10T05:06:39+00:00\",\"dateModified\":\"2025-12-10T05:25:49+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/experiment-results\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/experiment-results\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/experiment-results\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/wp-content\\\/uploads\\\/sites\\\/138\\\/2025\\\/12\\\/Screenshot-2025-12-10-at-12.18.45-AM.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/wp-content\\\/uploads\\\/sites\\\/138\\\/2025\\\/12\\\/Screenshot-2025-12-10-at-12.18.45-AM.png\",\"width\":832,\"height\":384},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/experiment-results\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Experiment Results\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/\",\"name\":\"2025 Team 16-1 Project\",\"description\":\"\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team16-1\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Experiment Results - 2025 Team 16-1 Project","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/","og_locale":"en_US","og_type":"article","og_title":"Experiment Results - 2025 Team 16-1 Project","og_description":"ComparisonsAs shown, our method achieves approximately 3\u00d7 faster inference compared to existing audio-driven and pose-driven baselines. In addition to speed, our approach produces higher-quality and more realistic results. Compared with audio-driven methods, our model not only maintains high generation quality but also substantially improves Sync-C and HKC. In particular, the lip synchronization confidence increases from &hellip; Continue reading \"Experiment Results\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/","og_site_name":"2025 Team 16-1 Project","article_modified_time":"2025-12-10T05:25:49+00:00","og_image":[{"width":832,"height":384,"url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.18.45-AM.png","type":"image\/png"}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"3 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/","name":"Experiment Results - 2025 Team 16-1 Project","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.18.45-AM.png","datePublished":"2025-12-10T05:06:39+00:00","dateModified":"2025-12-10T05:25:49+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.18.45-AM.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-content\/uploads\/sites\/138\/2025\/12\/Screenshot-2025-12-10-at-12.18.45-AM.png","width":832,"height":384},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/experiment-results\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/"},{"@type":"ListItem","position":2,"name":"Experiment Results"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/","name":"2025 Team 16-1 Project","description":"","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-json\/wp\/v2\/pages\/71","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-json\/wp\/v2\/users\/260"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-json\/wp\/v2\/comments?post=71"}],"version-history":[{"count":2,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-json\/wp\/v2\/pages\/71\/revisions"}],"predecessor-version":[{"id":79,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-json\/wp\/v2\/pages\/71\/revisions\/79"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team16-1\/wp-json\/wp\/v2\/media?parent=71"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}