{"id":133,"date":"2023-04-30T04:15:31","date_gmt":"2023-04-30T04:15:31","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/?page_id=133"},"modified":"2023-04-30T05:08:28","modified_gmt":"2023-04-30T05:08:28","slug":"data_uncertainty","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/","title":{"rendered":"Data Uncertainty"},"content":{"rendered":"\n<p><strong>Title<\/strong><\/p>\n\n\n\n<p>Visual Features Disentanglement into Interpretable Attributes Features using Prompts<\/p>\n\n\n\n<p><strong>Research Question: <\/strong><\/p>\n\n\n\n<p>How do we describe the distribution of image data?<\/p>\n\n\n\n<p><strong>Abstract<\/strong><\/p>\n\n\n\n<p>Deep Neural Networks are considered black-box mod- els as the feature embedding are often entangled and not human-interpretable. Though the feature embedding en- codes the information regarding understandable attributes (e.g., tail, head, wheel), it is distributed in the embedding and thus entangled. In this paper, we propose a simple method to disentangle dense image representations into a weighted combination of interpretable visual attributes by learning a simple linear projection in the joint space of a vision language model (VLM) in a data-driven fashion . By using images to weight attributes, our attributes are more semantically meaningful to the underlying data distribution than the generic curated attribute methods proposed in the literature. We demonstrate that our interpretable embed- ding performs competitively as compared to the dense em- bedding. Our discovered attributes act as better prompts under few shot settings, outperforming prompt-tuning methods (CoOp ) by 3% on 1 and 2 shot settings. Moreover, we also show a variety of potential downstream applications enabled by data-driven attributes, such as measuring domain shift, open vocabulary classifications and attribute- guided zero shot detection.<\/p>\n\n\n\n<p><strong>Contribution<\/strong><\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>We propose a simple yet efficient method to disentangle dense visual features to a weighted combination of interpretable attribute feature in language domain.&nbsp;<\/li>\n\n\n\n<li>We shows that our interpretable feature is a semantically meaningful yet a competitive feature in adapting visual concept by showing our superior performance on few shot evaluation.&nbsp;<\/li>\n\n\n\n<li>We also shows the potential of the attribute feature by showing a skew of downstream application like measuring distribution shrift, solving open-vocabulary classification and detection.<\/li>\n<\/ol>\n\n\n\n<p><strong>Motivation<\/strong><\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Feature embedding encodes the information regarding understandable attributes (e.g., tail, head, wheel) over data distribution, but it is distributed in the embedding and thus entangled.<\/li>\n\n\n\n<li>Semantic based disentangled feature can be interpretable data distribution measurement.&nbsp;<\/li>\n<\/ol>\n\n\n\n<figure class=\"wp-block-image size-large is-style-default\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"429\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1024x429.png\" alt=\"\" class=\"wp-image-147\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1024x429.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-300x126.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-768x322.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1536x644.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image.png 1682w\" sizes=\"auto, (max-width: 767px) 89vw, (max-width: 1000px) 54vw, (max-width: 1071px) 543px, 580px\" \/><figcaption class=\"wp-element-caption\">Visual Feature disentanglement<\/figcaption><\/figure>\n\n\n\n<p><strong>Pipeline<\/strong><\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Visual feature encoding and attribute text feature encoding with CLIP.<\/li>\n\n\n\n<li>Calculate pair-wise cosine similarity score between each attribute and class feature.<\/li>\n\n\n\n<li>Dataset adoption by performing linear probing over cosine similarity.&nbsp;<\/li>\n\n\n\n<li>Class-wise visual feature become a weighted combination of interpretable attribute feature.<\/li>\n<\/ol>\n\n\n\n<figure class=\"wp-block-image size-large is-style-default\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"503\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1-1024x503.png\" alt=\"\" class=\"wp-image-148\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1-1024x503.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1-300x147.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1-768x377.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1-1536x754.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1.png 1850w\" sizes=\"auto, (max-width: 767px) 89vw, (max-width: 1000px) 54vw, (max-width: 1071px) 543px, 580px\" \/><figcaption class=\"wp-element-caption\">Pipeline<\/figcaption><\/figure>\n\n\n\n<p><strong>Application<\/strong><\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li>Few shot generalization\n\n\n\n\n<ul class=\"wp-block-list\">\n<li>We aim to show that attributes generated in a data driven fashion from images collections could be a better visual descriptor compared to previous method which querying from LLM.<img decoding=\"async\" width=\"619px;\" height=\"204px;\" src=\"https:\/\/lh6.googleusercontent.com\/7YAV7SfUkyrSRnJyjd1Xwf_yZakKLMYvtdxs6hsOXUdgddmbKl7Y09pEiSb1Qv7TDCrgxwEAjxSgxu6iM0hPK8uryTRmCGkgD9gN2sTLMg3QBNFWLchfEx9mieeuN6XMzuDx6WLKAUIr7UHcKeWbJXmZPA=s2048\" alt=\"Table\n\nDescription automatically generated\"><\/li>\n<\/ul>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Our method outperform all previous finetuning method on 1 &amp; 2 &amp; 4 shot for a large margin. Compare to CoOp, which is a good baseline that use gradient propagation to learn the optimal embedding for prompting, but this will results in a non-interpretable feature. We shows the superiority of out method in terms of accuracy, interpretability and training time on extremely few shot setting.<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li>Quantifying distribution shift\n<ul class=\"wp-block-list\">\n<li>Quantifying distribution shift of ImageNet-sketch and -R(endition) dataset with respect to ImageNet1K dataset us- ing VAS-attributes (top-10 pos &amp; neg). We observe that color- less attributes become dominant and color attributes become subservient on ImageNet-sketch and art attributes become dominant on ImageNet-R. This demonstrates that we can capture the distribution shift in a human-interpretable way.\n<ul class=\"wp-block-list\">\n<li><img decoding=\"async\" width=\"461px;\" height=\"369px;\" src=\"https:\/\/lh3.googleusercontent.com\/POPeUoJKhNPkFl504hcGJJVpYKeJDaJtIMc2yxhnm-qfF-YwNKVc_Cv-aTNI9IKsRE7j4b_x2mE-GsbdgLTL4MryahQ_CKEIU5hCb4qk5IeXYrbYGgvo_9H38UxoTHKrlVFcFk3kioqm-r2g_4w6B5D27g=s2048\"><\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<\/li>\n\n\n\n<li>Open vocabulary object classification\n<ul class=\"wp-block-list\">\n<li>We show that it\u2019s possible to enable the classifier to gen- eralize onto novel concept(in this case, classify purple lemon) by swapping the order of our interpretable attributes feature.<\/li>\n\n\n\n<li><img decoding=\"async\" width=\"545px;\" height=\"418px;\" src=\"https:\/\/lh3.googleusercontent.com\/KU9o9rGbfHY_GCOxLTd02bJRfXzcFkU6Y7bPE-krbdJBf67MX_O8Fbyp-44nFHFlMUzMgW9N_RTUgKcd4UJzqRDl9ZqsA_tnxZRiisP9qWXNGJQjCInRE_q0YnAi1YRP0PtofuRPvCYhSebSo2O4y3D2oA=s2048\"><\/li>\n<\/ul>\n<\/li>\n\n\n\n<li>Attribute guided detection\n<ul class=\"wp-block-list\">\n<li>We repurpose CLIP RN50x64 for detection by computing the cosine similarity values between the Resnet embedding and our class-specific attribute list before the attention-pooling layer. This gives us a measure of which regions in the image correspond more to the input attributes, acting as a low-cost detector. On querying individually using our Hippopotamus and Nile Crocodile class attributes we can detect both classes respectively. These detections are obtained solely using our class-attributes which do not make use of the class-name.<\/li>\n<\/ul>\n<\/li>\n<\/ol>\n\n\n\n<p><img decoding=\"async\" width=\"764px;\" height=\"350px;\" src=\"https:\/\/lh6.googleusercontent.com\/sTdgXzy0pGxDydkLtd1ub-1a7BVLfgePIdrv9fCPgn-zxc2Nm7JT8DjIoT6zlXIDrrtOkQ2qWCDo4com2Hyy71zzecsQfxWr1oq0yXV4QZjTACGqw1T-H_bhn6psfT1GXE-JwYyMG4H8yg7d4NFxKqlbrw=s2048\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Title Visual Features Disentanglement into Interpretable Attributes Features using Prompts Research Question: How do we describe the distribution of image data? Abstract Deep Neural Networks are considered black-box mod- els as the feature embedding are often entangled and not human-interpretable. Though the feature embedding en- codes the information regarding understandable attributes (e.g., tail, head, wheel), &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Data Uncertainty&#8221;<\/span><\/a><\/p>\n","protected":false},"author":154,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-133","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Data Uncertainty - Uncertainty in Image classification<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Data Uncertainty - Uncertainty in Image classification\" \/>\n<meta property=\"og:description\" content=\"Title Visual Features Disentanglement into Interpretable Attributes Features using Prompts Research Question: How do we describe the distribution of image data? Abstract Deep Neural Networks are considered black-box mod- els as the feature embedding are often entangled and not human-interpretable. Though the feature embedding en- codes the information regarding understandable attributes (e.g., tail, head, wheel), &hellip; Continue reading &quot;Data Uncertainty&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/\" \/>\n<meta property=\"og:site_name\" content=\"Uncertainty in Image classification\" \/>\n<meta property=\"article:modified_time\" content=\"2023-04-30T05:08:28+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1024x429.png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"4 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/data_uncertainty\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/data_uncertainty\\\/\",\"name\":\"Data Uncertainty - Uncertainty in Image classification\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/data_uncertainty\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/data_uncertainty\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/wp-content\\\/uploads\\\/sites\\\/77\\\/2023\\\/04\\\/image-1024x429.png\",\"datePublished\":\"2023-04-30T04:15:31+00:00\",\"dateModified\":\"2023-04-30T05:08:28+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/data_uncertainty\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/data_uncertainty\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/data_uncertainty\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/wp-content\\\/uploads\\\/sites\\\/77\\\/2023\\\/04\\\/image.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/wp-content\\\/uploads\\\/sites\\\/77\\\/2023\\\/04\\\/image.png\",\"width\":1682,\"height\":705},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/data_uncertainty\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Data Uncertainty\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/\",\"name\":\"Uncertainty in Image classification\",\"description\":\"Students: Jia Shi | Advisors: Deva Ramanan (CMU), Shu Kong (CMU, TAMU), Francesco Ferroni (Argo AI), Arun Balajee Vasudevan (CMU) | Sponsor: Argo AI\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team6\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Data Uncertainty - Uncertainty in Image classification","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/","og_locale":"en_US","og_type":"article","og_title":"Data Uncertainty - Uncertainty in Image classification","og_description":"Title Visual Features Disentanglement into Interpretable Attributes Features using Prompts Research Question: How do we describe the distribution of image data? Abstract Deep Neural Networks are considered black-box mod- els as the feature embedding are often entangled and not human-interpretable. Though the feature embedding en- codes the information regarding understandable attributes (e.g., tail, head, wheel), &hellip; Continue reading \"Data Uncertainty\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/","og_site_name":"Uncertainty in Image classification","article_modified_time":"2023-04-30T05:08:28+00:00","og_image":[{"url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1024x429.png","type":"","width":"","height":""}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/","name":"Data Uncertainty - Uncertainty in Image classification","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image-1024x429.png","datePublished":"2023-04-30T04:15:31+00:00","dateModified":"2023-04-30T05:08:28+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-content\/uploads\/sites\/77\/2023\/04\/image.png","width":1682,"height":705},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/data_uncertainty\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/"},{"@type":"ListItem","position":2,"name":"Data Uncertainty"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/","name":"Uncertainty in Image classification","description":"Students: Jia Shi | Advisors: Deva Ramanan (CMU), Shu Kong (CMU, TAMU), Francesco Ferroni (Argo AI), Arun Balajee Vasudevan (CMU) | Sponsor: Argo AI","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-json\/wp\/v2\/pages\/133","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-json\/wp\/v2\/users\/154"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-json\/wp\/v2\/comments?post=133"}],"version-history":[{"count":5,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-json\/wp\/v2\/pages\/133\/revisions"}],"predecessor-version":[{"id":151,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-json\/wp\/v2\/pages\/133\/revisions\/151"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team6\/wp-json\/wp\/v2\/media?parent=133"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}