{"id":20,"date":"2022-05-01T04:39:08","date_gmt":"2022-05-01T04:39:08","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/?page_id=20"},"modified":"2022-12-19T19:05:13","modified_gmt":"2022-12-19T19:05:13","slug":"pose-estimation","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/","title":{"rendered":"Pose Estimation"},"content":{"rendered":"\n<h3 class=\"wp-block-heading\"><strong>Problem Statement<\/strong><\/h3>\n\n\n\n<p>Given a set of cameras and their known intrinsic parameters, estimate the 6DoF pose ( i.e., Calculate Rotation and Translation parameters) of each of the cameras present in the setting.<\/p>\n\n\n\n<figure class=\"wp-block-image is-style-default\"><img decoding=\"async\" src=\"https:\/\/dovkatz.files.wordpress.com\/2018\/06\/sfm-e1528657740297.jpg?w=640\" alt=\"3D Sensing \u2013 Dov Katz: computer vision\" \/><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Solution<\/strong><\/h3>\n\n\n\n<p>This is a well-researched problem and can be solved from Structure from Motion(SfM) and a Deep learning-based approach. But\u00a0<strong>SfM<\/strong>\u00a0approach performs better than deep learning-based. Hence Structure from motion approach is used to obtain the Rotation and Translation parameters of the cameras. Here is the proposed pipeline of the work based on Incremental SfM<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"197\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-5-1024x197.png\" alt=\"\" class=\"wp-image-147\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-5-1024x197.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-5-300x58.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-5-768x147.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-5-1536x295.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-5.png 1604w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption>Image credits: colmap.github.io<\/figcaption><\/figure>\n\n\n\n<h4 class=\"wp-block-heading\">Correspondence search  and matching<\/h4>\n\n\n\n<figure class=\"wp-block-image size-large is-style-default\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"160\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-3-1024x160.png\" alt=\"\" class=\"wp-image-97\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-3-1024x160.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-3-300x47.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-3-768x120.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-3.png 1515w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption>Proposed Pipeline for correspondence search  and matching<\/figcaption><\/figure>\n\n\n\n<h5 class=\"wp-block-heading\">Keypoint Detection and Description<\/h5>\n\n\n\n<p>Identifying keypoints and describing them is an ill posed task as there is no correct solution for the porblem. A keypoint detector and descriptor should find interesting points in the images and describe them in a manner which would be invariant to viewpoint, illumination etc. I have chosen to use SuperPoint model for these tasks. It is a self supervised framework for training interest point detectors and descriptors suitable for a large number of multiple-view geometry problems in computer vision.<\/p>\n\n\n\n<h5 class=\"wp-block-heading\">Feature Matching<\/h5>\n\n\n\n<p>It is a task of matching keypoints from one image onto another image capturing same part of the image. This helps us finding correspondences in the images and removes outliers. I have opted SuperGlue model for this. It is a graphical neural network approach to match keypoints. One the the interesting things of this method is to take whole image as input for match as opposed to part of image which helps it avoid mismatching points of similar texture.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large is-style-default\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"380\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-5-1024x380.png\" alt=\"\" class=\"wp-image-99\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-5-1024x380.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-5-300x111.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-5-768x285.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/05\/image-5.png 1214w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption>Sample Result For Keypoint Detection and Matching<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"390\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-6-1024x390.png\" alt=\"\" class=\"wp-image-148\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-6-1024x390.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-6-300x114.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-6-768x293.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-6-1536x585.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-6.png 1596w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption>SuperGlue keypoint matching result on stronger threshold<\/figcaption><\/figure>\n\n\n\n<h4 class=\"wp-block-heading\">Pose Estimation<br><\/h4>\n\n\n\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<p>This solution approach is inspired by classical multi-view  geometry. Let us consider N set of cameras. Initially, we will select a pair of cameras and undistort their images. Next we would get keypoint correspondences and compute Fundamental matrix between them using a standard 8point algorithm. Later we obtain Essential matrix from Fundamental matrix and decompose it into Rotational and Translation parameters. Furthermore we can obtain 3D points using triangulation and perform local  bundle adjustment to optimize the parameters. For each successive camera, we try to find 2D-3D correspondence pairs in the image with the help of previous images and their projected 3D points and solve PnP problem to obtain its camera pose.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Evaluation Metric<\/strong><\/h3>\n<\/div><\/div>\n\n\n\n<p> For evaluation, I have taken Reprojection error as metric. Reprojection error as the name suggests an error on projecting one key point from an image to another image.<br><\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><strong>Result<\/strong><\/h3>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"457\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-7-1024x457.png\" alt=\"\" class=\"wp-image-149\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-7-1024x457.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-7-300x134.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-7-768x343.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-7-1536x686.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-content\/uploads\/sites\/67\/2022\/12\/image-7.png 1819w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption>Visualization of sparse 3D scene and camera position<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-table is-style-regular\"><table><tbody><tr><td><strong>Stats<\/strong><\/td><td><strong>SuperPoint + SuperGlue<\/strong><\/td><\/tr><tr><td>Number of 3D points<\/td><td> 747<\/td><\/tr><tr><td>Mean observation per image<\/td><td>589<\/td><\/tr><tr><td>Mean Reprojection error (in px)<\/td><td>0.36<\/td><\/tr><\/tbody><\/table><\/figure>\n","protected":false},"excerpt":{"rendered":"<p>Problem Statement Given a set of cameras and their known intrinsic parameters, estimate the 6DoF pose ( i.e., Calculate Rotation and Translation parameters) of each of the cameras present in the setting. Solution This is a well-researched problem and can be solved from Structure from Motion(SfM) and a Deep learning-based approach. But\u00a0SfM\u00a0approach performs better than &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Pose Estimation&#8221;<\/span><\/a><\/p>\n","protected":false},"author":123,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-20","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Pose Estimation - Multi-Camera Pose Estimation<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Pose Estimation - Multi-Camera Pose Estimation\" \/>\n<meta property=\"og:description\" content=\"Problem Statement Given a set of cameras and their known intrinsic parameters, estimate the 6DoF pose ( i.e., Calculate Rotation and Translation parameters) of each of the cameras present in the setting. Solution This is a well-researched problem and can be solved from Structure from Motion(SfM) and a Deep learning-based approach. But\u00a0SfM\u00a0approach performs better than &hellip; Continue reading &quot;Pose Estimation&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/\" \/>\n<meta property=\"og:site_name\" content=\"Multi-Camera Pose Estimation\" \/>\n<meta property=\"article:modified_time\" content=\"2022-12-19T19:05:13+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/dovkatz.files.wordpress.com\/2018\/06\/sfm-e1528657740297.jpg?w=640\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"4 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/pose-estimation\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/pose-estimation\\\/\",\"name\":\"Pose Estimation - Multi-Camera Pose Estimation\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/pose-estimation\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/pose-estimation\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/dovkatz.files.wordpress.com\\\/2018\\\/06\\\/sfm-e1528657740297.jpg?w=640\",\"datePublished\":\"2022-05-01T04:39:08+00:00\",\"dateModified\":\"2022-12-19T19:05:13+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/pose-estimation\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/pose-estimation\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/pose-estimation\\\/#primaryimage\",\"url\":\"https:\\\/\\\/dovkatz.files.wordpress.com\\\/2018\\\/06\\\/sfm-e1528657740297.jpg?w=640\",\"contentUrl\":\"https:\\\/\\\/dovkatz.files.wordpress.com\\\/2018\\\/06\\\/sfm-e1528657740297.jpg?w=640\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/pose-estimation\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Pose Estimation\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/\",\"name\":\"Multi-Camera Pose Estimation\",\"description\":\"Student: Aditya Ghuge | Advisor: Fernando De La Torre | Sponsor: Komatsu\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2022team12\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Pose Estimation - Multi-Camera Pose Estimation","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/","og_locale":"en_US","og_type":"article","og_title":"Pose Estimation - Multi-Camera Pose Estimation","og_description":"Problem Statement Given a set of cameras and their known intrinsic parameters, estimate the 6DoF pose ( i.e., Calculate Rotation and Translation parameters) of each of the cameras present in the setting. Solution This is a well-researched problem and can be solved from Structure from Motion(SfM) and a Deep learning-based approach. But\u00a0SfM\u00a0approach performs better than &hellip; Continue reading \"Pose Estimation\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/","og_site_name":"Multi-Camera Pose Estimation","article_modified_time":"2022-12-19T19:05:13+00:00","og_image":[{"url":"https:\/\/dovkatz.files.wordpress.com\/2018\/06\/sfm-e1528657740297.jpg?w=640","type":"","width":"","height":""}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/","name":"Pose Estimation - Multi-Camera Pose Estimation","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/#primaryimage"},"thumbnailUrl":"https:\/\/dovkatz.files.wordpress.com\/2018\/06\/sfm-e1528657740297.jpg?w=640","datePublished":"2022-05-01T04:39:08+00:00","dateModified":"2022-12-19T19:05:13+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/#primaryimage","url":"https:\/\/dovkatz.files.wordpress.com\/2018\/06\/sfm-e1528657740297.jpg?w=640","contentUrl":"https:\/\/dovkatz.files.wordpress.com\/2018\/06\/sfm-e1528657740297.jpg?w=640"},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/pose-estimation\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/"},{"@type":"ListItem","position":2,"name":"Pose Estimation"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/","name":"Multi-Camera Pose Estimation","description":"Student: Aditya Ghuge | Advisor: Fernando De La Torre | Sponsor: Komatsu","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-json\/wp\/v2\/pages\/20","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-json\/wp\/v2\/users\/123"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-json\/wp\/v2\/comments?post=20"}],"version-history":[{"count":5,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-json\/wp\/v2\/pages\/20\/revisions"}],"predecessor-version":[{"id":150,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-json\/wp\/v2\/pages\/20\/revisions\/150"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2022team12\/wp-json\/wp\/v2\/media?parent=20"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}