{"id":3,"date":"2025-05-09T20:59:44","date_gmt":"2025-05-09T20:59:44","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/?p=3"},"modified":"2025-05-10T00:41:02","modified_gmt":"2025-05-10T00:41:02","slug":"temporally-hierarchical-scene-graph-generation-for-video-question-answering","status":"publish","type":"post","link":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/","title":{"rendered":"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b"},"content":{"rendered":"\n<p>2025 MSCV Capstone project of Daniel Yang, Tianzhi Li<br>Supervised by <a href=\"https:\/\/www.ri.cmu.edu\/ri-faculty\/katia-sycara\/\" target=\"_blank\" rel=\"noreferrer noopener\">Katia Sycara<\/a>, Yaqi Xie<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Motivation<\/h2>\n\n\n\n<p class=\"has-small-font-size\"><strong>Video Question Answering (VideoQA)<\/strong> requires reasoning over actions that occur at drastically different timescales \u2014 from short, atomic gestures like picking up an apple to long-term goals like shopping in the supermarket. Traditional vision-language models struggle to capture such temporal diversity, particularly in long videos where long, overarching intentions and fine-grained actions coexist.<br>We propose a <strong>temporally hierarchical scene graph<\/strong> representation to address this. Scene graphs encode video content as compact &lt;Subject &#8211; Relation &#8211; Object&gt; triplets, enabling structured and explainable reasoning. By building scene graphs at multiple temporal resolutions \u2014 from frame-level interactions to high-level goals \u2014 we can preserve both fine-grained actions and overarching intentions.<br>This approach offers a scalable, interpretable, and resource-efficient method for long-form video understanding, making VideoQA more tractable across diverse tasks and durations.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"649\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2-1024x649.png\" alt=\"\" class=\"wp-image-9\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2-1024x649.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2-300x190.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2-768x487.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2.png 1110w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\">An example of an image scene graph that captures the semantics of a scene.<\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Datasets<\/h2>\n\n\n\n<p class=\"has-medium-font-size\"><a href=\"https:\/\/egoschema.github.io\/\" target=\"_blank\" rel=\"noreferrer noopener\">EgoSchema<\/a><\/p>\n\n\n\n<p class=\"has-small-font-size\">The EgoSchema benchmark contains over 5000 very long-form video language understanding questions spanning over 250 hours of real, diverse, and high-quality egocentric video data. Many videos feature complex scenes with cluttered rooms, where we could take full potential of scene graphs.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"721\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-5-1024x721.png\" alt=\"\" class=\"wp-image-12\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-5-1024x721.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-5-300x211.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-5-768x541.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-5.png 1286w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\">EgoSchema dataset QA example<\/figcaption><\/figure>\n\n\n\n<p class=\"has-medium-font-size\"><a href=\"https:\/\/github.com\/doc-doc\/NExT-QA\" target=\"_blank\" rel=\"noreferrer noopener\">NExT-QA<\/a><\/p>\n\n\n\n<p class=\"has-small-font-size\">NExT-QA is a VideoQA benchmark targeting the explanation of video content. It challenges QA models to reason about the causal and temporal actions and understand the rich object interactions in daily activities. NExT-QA contains 5,440 videos and about 52K manually annotated question-answer pairs, grouped into causal, temporal, and descriptive questions. The videos are about 35 seconds long and focus on interactions between people and sometimes pets. <\/p>\n\n\n\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"968\" height=\"706\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-3.png\" alt=\"\" class=\"wp-image-10\" style=\"width:602px;height:auto\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-3.png 968w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-3-300x219.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-3-768x560.png 768w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\">NExT-QA examples and question category statistics<\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Proposed Method<\/h2>\n\n\n\n<p class=\"has-small-font-size\">We plan to simultaneously generate a hierarchical structure of scene graphs and utilize them to tackle VideoQA tasks. <\/p>\n\n\n\n<p class=\"has-small-font-size\">Since we don&#8217;t have access to higher-level scene graphs beyond frame-by-frame image scene graphs, we have to construct the higher-level scene graphs in a self-supervised manner.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"938\" height=\"272\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-4.png\" alt=\"\" class=\"wp-image-11\" style=\"width:439px;height:auto\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-4.png 938w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-4-300x87.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-4-768x223.png 768w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\">The key source of supervision for scene graph encoding: Variational AutoEncoder (VAE)<\/figcaption><\/figure>\n\n\n\n<p class=\"has-small-font-size\">We outlined the steps to generate high-level scene graphs that express short clips from image-level scene graphs.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"301\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-1-1024x301.png\" alt=\"\" class=\"wp-image-8\" style=\"width:647px;height:auto\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-1-1024x301.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-1-300x88.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-1-768x226.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-1-1536x452.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-1.png 1720w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\">Architecture overview of our proposed temporally hierarchical scene graph generator.<\/figcaption><\/figure>\n\n\n\n<ul class=\"wp-block-list\">\n<li class=\"has-small-font-size\"><strong>Dense low-level scene graph prediction from video frames<\/strong>\n<ul class=\"wp-block-list\">\n<li>Densely predict scene graph for each frame with off-the-shelf scene graph generator; collapse identical and continuous scene graphs.<\/li>\n\n\n\n<li>Each node has a node embedding, initialized with image embedding &amp; word embedding.<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li class=\"has-small-font-size\"><strong>Self-supervised scene graph encoding<\/strong>\n<ul class=\"wp-block-list\">\n<li>Given the lack of supervision for scene graph encoding, we adopt a Variational Autoencoder (VAE) to learn embeddings in a self-supervised manner.<\/li>\n\n\n\n<li>The encoder maps the scene graph into a continuous latent vector, enforcing semantic smoothness across similar graphs.<\/li>\n\n\n\n<li>We also adopt a contrastive loss to align the latent vector of temporally adjacent scene graphs.<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li class=\"has-small-font-size\"><strong>Local action segmentation &amp; scene graph reconstruction<\/strong>\n<ul class=\"wp-block-list\">\n<li>Given the sequence of latent vectors from the scene graph encoder, we segment the video into coherent action segments.<\/li>\n\n\n\n<li>Aggregate latent vectors within segments via self-attention to obtain high-level segment embeddings.<\/li>\n\n\n\n<li>Utilize the VAE decoder to reconstruct segment-level scene graphs, forming a higher-level abstraction.<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li class=\"has-small-font-size\"><strong>Video QA end-to-end finetuning<\/strong>\n<ul class=\"wp-block-list\">\n<li>Tokenize higher-level scene graphs and input them to an LLM for question answering.<\/li>\n\n\n\n<li>Finetune the entire system end-to-end, with LoRA-based updates on the LLM and full updates on upstream modules<\/li>\n\n\n\n<li class=\"has-small-font-size\">For subsequent QAs, we should be able to discard the original video and answer questions with our hierarchical scene graph structure.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">References<\/h2>\n\n\n\n<p class=\"has-small-font-size\">[1] Mangalam, K., Akshulakov, R., &amp; Malik, J. (2023). <em>EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding<\/em>. arXiv:2308.09126. <a class=\"\" href=\"https:\/\/arxiv.org\/abs\/2308.09126\">https:\/\/arxiv.org\/abs\/2308.09126<\/a><\/p>\n\n\n\n<p class=\"has-small-font-size\">[2] Xiao, Junbin, et al. &#8220;Next-qa: Next phase of question-answering to explaining temporal actions.&#8221; <em>Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition<\/em>. 2021.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[3] Wang, Guan, et al. &#8220;OED: towards one-stage end-to-end dynamic scene graph generation.&#8221; <em>Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition<\/em>. 2024.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[4] Islam, Md Mohaiminul, et al. &#8220;Video recap: Recursive captioning of hour-long videos.&#8221; <em>Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition<\/em>. 2024.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[5] Liang, Weixin, Yanhao Jiang, and Zixuan Liu. &#8220;GraghVQA: Language-guided graph neural networks for graph-based visual question answering.&#8221; <em>arXiv preprint arXiv:2104.10283<\/em> (2021).[6] Nag, Sayak, et al. &#8220;Unbiased scene graph generation in videos.&#8221; <em>Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition<\/em>. 2023.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>2025 MSCV Capstone project of Daniel Yang, Tianzhi LiSupervised by Katia Sycara, Yaqi Xie Motivation Video Question Answering (VideoQA) requires reasoning over actions that occur at drastically different timescales \u2014 from short, atomic gestures like picking up an apple to long-term goals like shopping in the supermarket. Traditional vision-language models struggle to capture such temporal &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b&#8221;<\/span><\/a><\/p>\n","protected":false},"author":239,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-3","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b - Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b - Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b\" \/>\n<meta property=\"og:description\" content=\"2025 MSCV Capstone project of Daniel Yang, Tianzhi LiSupervised by Katia Sycara, Yaqi Xie Motivation Video Question Answering (VideoQA) requires reasoning over actions that occur at drastically different timescales \u2014 from short, atomic gestures like picking up an apple to long-term goals like shopping in the supermarket. Traditional vision-language models struggle to capture such temporal &hellip; Continue reading &quot;Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/\" \/>\n<meta property=\"og:site_name\" content=\"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b\" \/>\n<meta property=\"article:published_time\" content=\"2025-05-09T20:59:44+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2025-05-10T00:41:02+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2-1024x649.png\" \/>\n<meta name=\"author\" content=\"danielya\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"danielya\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"5 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/\"},\"author\":{\"name\":\"danielya\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/#\\\/schema\\\/person\\\/20aa2a24dc5e112653ad2b242a56b7e7\"},\"headline\":\"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b\",\"datePublished\":\"2025-05-09T20:59:44+00:00\",\"dateModified\":\"2025-05-10T00:41:02+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/\"},\"wordCount\":709,\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/wp-content\\\/uploads\\\/sites\\\/124\\\/2025\\\/05\\\/image-2-1024x649.png\",\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/\",\"name\":\"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b - Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/wp-content\\\/uploads\\\/sites\\\/124\\\/2025\\\/05\\\/image-2-1024x649.png\",\"datePublished\":\"2025-05-09T20:59:44+00:00\",\"dateModified\":\"2025-05-10T00:41:02+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/#\\\/schema\\\/person\\\/20aa2a24dc5e112653ad2b242a56b7e7\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/wp-content\\\/uploads\\\/sites\\\/124\\\/2025\\\/05\\\/image-2.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/wp-content\\\/uploads\\\/sites\\\/124\\\/2025\\\/05\\\/image-2.png\",\"width\":1110,\"height\":704},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/2025\\\/05\\\/09\\\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/\",\"name\":\"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b\",\"description\":\"\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/#\\\/schema\\\/person\\\/20aa2a24dc5e112653ad2b242a56b7e7\",\"name\":\"danielya\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/0afd82786e7c8c454f8922102d1946e909c9f987a7839234adc668771a929492?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/0afd82786e7c8c454f8922102d1946e909c9f987a7839234adc668771a929492?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/0afd82786e7c8c454f8922102d1946e909c9f987a7839234adc668771a929492?s=96&d=mm&r=g\",\"caption\":\"danielya\"},\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2025team17\\\/author\\\/danielya\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b - Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/","og_locale":"en_US","og_type":"article","og_title":"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b - Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b","og_description":"2025 MSCV Capstone project of Daniel Yang, Tianzhi LiSupervised by Katia Sycara, Yaqi Xie Motivation Video Question Answering (VideoQA) requires reasoning over actions that occur at drastically different timescales \u2014 from short, atomic gestures like picking up an apple to long-term goals like shopping in the supermarket. Traditional vision-language models struggle to capture such temporal &hellip; Continue reading \"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/","og_site_name":"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b","article_published_time":"2025-05-09T20:59:44+00:00","article_modified_time":"2025-05-10T00:41:02+00:00","og_image":[{"url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2-1024x649.png","type":"","width":"","height":""}],"author":"danielya","twitter_card":"summary_large_image","twitter_misc":{"Written by":"danielya","Est. reading time":"5 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/#article","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/"},"author":{"name":"danielya","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/#\/schema\/person\/20aa2a24dc5e112653ad2b242a56b7e7"},"headline":"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b","datePublished":"2025-05-09T20:59:44+00:00","dateModified":"2025-05-10T00:41:02+00:00","mainEntityOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/"},"wordCount":709,"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2-1024x649.png","inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/","name":"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b - Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2-1024x649.png","datePublished":"2025-05-09T20:59:44+00:00","dateModified":"2025-05-10T00:41:02+00:00","author":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/#\/schema\/person\/20aa2a24dc5e112653ad2b242a56b7e7"},"breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-content\/uploads\/sites\/124\/2025\/05\/image-2.png","width":1110,"height":704},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/2025\/05\/09\/temporally-hierarchical-scene-graph-generation-for-video-question-answering\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/"},{"@type":"ListItem","position":2,"name":"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/","name":"Temporally Hierarchical Scene Graph Generation for Video Question Answering\u200b","description":"","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/#\/schema\/person\/20aa2a24dc5e112653ad2b242a56b7e7","name":"danielya","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/0afd82786e7c8c454f8922102d1946e909c9f987a7839234adc668771a929492?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/0afd82786e7c8c454f8922102d1946e909c9f987a7839234adc668771a929492?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/0afd82786e7c8c454f8922102d1946e909c9f987a7839234adc668771a929492?s=96&d=mm&r=g","caption":"danielya"},"url":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/author\/danielya\/"}]}},"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/posts\/3","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/users\/239"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/comments?post=3"}],"version-history":[{"count":3,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/posts\/3\/revisions"}],"predecessor-version":[{"id":15,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/posts\/3\/revisions\/15"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/media?parent=3"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/categories?post=3"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2025team17\/wp-json\/wp\/v2\/tags?post=3"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}