{"id":46,"date":"2023-05-04T16:31:22","date_gmt":"2023-05-04T16:31:22","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/?page_id=46"},"modified":"2023-12-18T02:50:21","modified_gmt":"2023-12-18T02:50:21","slug":"method","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/","title":{"rendered":"Method"},"content":{"rendered":"\n<h2 class=\"wp-block-heading\">Overview<\/h2>\n\n\n\n<p>Unlike most previous works, we associate information across cameras at the detection stage to get an accurate top-down view detection results. Next, we track and project points back to each camera view. A single-view detector will be used to further refine the bounding boxes. We also extracts appearance feature for more robust association.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"180\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM-1024x180.png\" alt=\"\" class=\"wp-image-177\" style=\"width:791px;height:auto\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM-1024x180.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM-300x53.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM-768x135.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM-1536x269.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM.png 1870w\" sizes=\"auto, (max-width: 767px) 89vw, (max-width: 1000px) 54vw, (max-width: 1071px) 543px, 580px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Multi-View Detection<\/h2>\n\n\n\n<p>Based on recent research in BEV perception or 3D detection [1] [2], we build the following architecture. The model runs a modified ResNet backbone on each views and then projects the extracted features into the top-down view using camera calibration data. Concatenating these projected features aggregates information across camera views. For the spatial aggregation module, there are two options: large-kernel convolution and deformable transformer. The latter option allows a larger receptive field and is proved to perform better in our case. <\/p>\n\n\n\n<p>A single view detection head is also trained and its bounding box regression loss is used as an auxiliary loss term of feature extractor.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"667\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/05\/Screen-Shot-2023-05-04-at-1.17.37-PM-1024x667.png\" alt=\"\" class=\"wp-image-85\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/05\/Screen-Shot-2023-05-04-at-1.17.37-PM-1024x667.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/05\/Screen-Shot-2023-05-04-at-1.17.37-PM-300x195.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/05\/Screen-Shot-2023-05-04-at-1.17.37-PM-768x500.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/05\/Screen-Shot-2023-05-04-at-1.17.37-PM.png 1524w\" sizes=\"auto, (max-width: 767px) 89vw, (max-width: 1000px) 54vw, (max-width: 1071px) 543px, 580px\" \/><figcaption class=\"wp-element-caption\">Diagram from https:\/\/arxiv.org\/abs\/2007.07247<br><\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Projection and Refinement<\/h2>\n\n\n\n<p>We assume a fixed human body size which allows us to project 3D detection results back to camera views. The projected bounding boxes can have an inaccurate shape but can be refined by matching them with results from normal single-view detectors. Another benefit of this refinement is that we can identify and remove false positives when some boxes are not matched in any of the views.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"878\" height=\"540\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/05\/Screen-Shot-2023-05-04-at-1.52.52-PM.png\" alt=\"\" class=\"wp-image-123\" style=\"width:619px;height:auto\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/05\/Screen-Shot-2023-05-04-at-1.52.52-PM.png 878w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/05\/Screen-Shot-2023-05-04-at-1.52.52-PM-300x185.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/05\/Screen-Shot-2023-05-04-at-1.52.52-PM-768x472.png 768w\" sizes=\"auto, (max-width: 767px) 89vw, (max-width: 1000px) 54vw, (max-width: 1071px) 543px, 580px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Tracking<\/h2>\n\n\n\n<p>We start with SORT [3] with an adjustment of its Kalman filter to track points in the top-down view, and later upgrade to a DeepSORT-like [5] tracking algorithm. Using refined bounding boxes, we can incorporate appearance feature at the association stage which allows &#8220;remembering&#8221; a target even if we lose track of him\/her for a short period. We explore minimum cosine distance and averaging as two ways of aggregating appearance feature in the multi-view setting. <\/p>\n\n\n\n<p>More specifically, we filter out features from occluded targets (invisible to single-view detector) and use, for example,  average across views to represent its multi-view appearance. Finally, the min cosine distance between feature vector from new detections with those from tracking history is used as the appearance distance in matching.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"602\" height=\"330\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.32.38-PM.png\" alt=\"\" class=\"wp-image-178\" style=\"aspect-ratio:1.9647577092511013;width:466px;height:auto\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.32.38-PM.png 602w, https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.32.38-PM-300x164.png 300w\" sizes=\"auto, (max-width: 602px) 100vw, 602px\" \/><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">Experiment<\/h2>\n\n\n\n<p>We train and evaluate our model on MMPTRACK dataset [4], which is a large-scale video dataset for multi-camera multi-object tracking. The dataset has ~5 hour videos for training and 1.5 hour videos for validation. Annotations include per-frame bounding boxes, corresponding person IDs and camera calibration data. The dataset poses challenges like cluttered and crowded environments, varying human poses and appearances to our tracking system.<\/p>\n\n\n\n<p>We mostly focus on the retail environment and fine-tuned the pre-trained MVDeTr dataset for 10 epochs sampling 1 image every 5 images (for nearby frames in videos look alike). As  images from all cameras need to be processed, we reduce the size of images by a factor of 4 to reduce the memory overhead. The world coordinate is voxelized with a size of 20mm \u00d7 20mm \u00d7 20mm and the ground plane is further reduced by a factor of 2. For the single view detector, we fine-tune YOLOv8-small for 20 epochs predicting Person class. Finally, an off-the-shelf Market1501 pre-trained OSNet is used to extract appearance features.<\/p>\n\n\n\n<p>For quantitative evaluation, we mainly uses MOTA and IDF1 scores. We set up matching threshold as IoU&gt;0.5 for camera view and pixel distance &lt;25  in top-down view, following rules used in the MMPTRACK challenge. For updated results, please visit our <a href=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/fall-2023\/\">2023 Fall<\/a> page.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Reference<\/h2>\n\n\n\n<p class=\"has-small-font-size\">[1] Hou, Yunzhong, Liang Zheng, and Stephen Gould. &#8220;Multiview detection with feature perspective transformation.&#8221;&nbsp;<em>Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part VII 16<\/em>. Springer International Publishing, 2020.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[2] Hou, Yunzhong, and Liang Zheng. &#8220;Multiview detection with shadow transformer (and view-coherent data augmentation).&#8221;&nbsp;<em>Proceedings of the 29th ACM International Conference on Multimedia<\/em>. 2021.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[3] Bewley, Alex, et al. &#8220;Simple online and realtime tracking.&#8221;&nbsp;<em>2016 IEEE international conference on image processing (ICIP)<\/em>. IEEE, 2016.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[4] Han, Xiaotian, et al. &#8220;Mmptrack: Large-scale densely annotated multi-camera multiple people tracking benchmark.&#8221;&nbsp;<em>Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision<\/em>. 2023.<\/p>\n\n\n\n<p class=\"has-small-font-size\">[5] Wojke, Nicolai, Alex Bewley, and Dietrich Paulus. &#8220;Simple online and realtime tracking with a deep association metric.&#8221;&nbsp;<em>2017 IEEE international conference on image processing (ICIP)<\/em>. IEEE, 2017.<\/p>\n\n\n\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Overview Unlike most previous works, we associate information across cameras at the detection stage to get an accurate top-down view detection results. Next, we track and project points back to each camera view. A single-view detector will be used to further refine the bounding boxes. We also extracts appearance feature for more robust association. Multi-View &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Method&#8221;<\/span><\/a><\/p>\n","protected":false},"author":167,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-46","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Method - Multi-Camera Multi-People Tracking<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Method - Multi-Camera Multi-People Tracking\" \/>\n<meta property=\"og:description\" content=\"Overview Unlike most previous works, we associate information across cameras at the detection stage to get an accurate top-down view detection results. Next, we track and project points back to each camera view. A single-view detector will be used to further refine the bounding boxes. We also extracts appearance feature for more robust association. Multi-View &hellip; Continue reading &quot;Method&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/\" \/>\n<meta property=\"og:site_name\" content=\"Multi-Camera Multi-People Tracking\" \/>\n<meta property=\"article:modified_time\" content=\"2023-12-18T02:50:21+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM-1024x180.png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"4 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/method\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/method\\\/\",\"name\":\"Method - Multi-Camera Multi-People Tracking\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/method\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/method\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/wp-content\\\/uploads\\\/sites\\\/84\\\/2023\\\/12\\\/Screen-Shot-2023-12-16-at-9.23.53-PM-1024x180.png\",\"datePublished\":\"2023-05-04T16:31:22+00:00\",\"dateModified\":\"2023-12-18T02:50:21+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/method\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/method\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/method\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/wp-content\\\/uploads\\\/sites\\\/84\\\/2023\\\/12\\\/Screen-Shot-2023-12-16-at-9.23.53-PM.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/wp-content\\\/uploads\\\/sites\\\/84\\\/2023\\\/12\\\/Screen-Shot-2023-12-16-at-9.23.53-PM.png\",\"width\":1870,\"height\":328},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/method\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Method\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/\",\"name\":\"Multi-Camera Multi-People Tracking\",\"description\":\"CMU MSCV &#039;23 Capstone Project sponsored by Centific\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/f23team7\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Method - Multi-Camera Multi-People Tracking","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/","og_locale":"en_US","og_type":"article","og_title":"Method - Multi-Camera Multi-People Tracking","og_description":"Overview Unlike most previous works, we associate information across cameras at the detection stage to get an accurate top-down view detection results. Next, we track and project points back to each camera view. A single-view detector will be used to further refine the bounding boxes. We also extracts appearance feature for more robust association. Multi-View &hellip; Continue reading \"Method\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/","og_site_name":"Multi-Camera Multi-People Tracking","article_modified_time":"2023-12-18T02:50:21+00:00","og_image":[{"url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM-1024x180.png","type":"","width":"","height":""}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/","name":"Method - Multi-Camera Multi-People Tracking","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM-1024x180.png","datePublished":"2023-05-04T16:31:22+00:00","dateModified":"2023-12-18T02:50:21+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-content\/uploads\/sites\/84\/2023\/12\/Screen-Shot-2023-12-16-at-9.23.53-PM.png","width":1870,"height":328},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/method\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/"},{"@type":"ListItem","position":2,"name":"Method"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/","name":"Multi-Camera Multi-People Tracking","description":"CMU MSCV &#039;23 Capstone Project sponsored by Centific","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-json\/wp\/v2\/pages\/46","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-json\/wp\/v2\/users\/167"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-json\/wp\/v2\/comments?post=46"}],"version-history":[{"count":11,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-json\/wp\/v2\/pages\/46\/revisions"}],"predecessor-version":[{"id":194,"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-json\/wp\/v2\/pages\/46\/revisions\/194"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/f23team7\/wp-json\/wp\/v2\/media?parent=46"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}