{"id":42,"date":"2022-12-20T00:17:52","date_gmt":"2022-12-20T00:17:52","guid":{"rendered":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/?page_id=42"},"modified":"2023-05-08T04:45:06","modified_gmt":"2023-05-08T04:45:06","slug":"introduction","status":"publish","type":"page","link":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/","title":{"rendered":"Experiments"},"content":{"rendered":"\n<p>In this section, we validate our method (ITI-Gen) for inclusive text-to-image generation on various attributes and scenarios.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Single Binary &amp; Multiple<\/strong> <strong>Attributes<\/strong><\/h2>\n\n\n\n<p>To demonstrate the capability of our method to sample images with a variety of face attributes, we construct 40 distinct reference image sets based on attributes from CelebA [1]. Also, given multiple reference image sets (each captures the marginal distribution for an attribute), ITI-GEN can generate diverse images across any category combination of the attributes.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large is-resized\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png\" alt=\"\" class=\"wp-image-248\" width=\"674\" height=\"298\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-300x133.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-768x341.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-1536x681.png 1536w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM.png 1628w\" sizes=\"auto, (max-width: 674px) 100vw, 674px\" \/><figcaption class=\"wp-element-caption\"><em>Figure 1. <\/em>Comparison with baseline methods with single attribute and multiple attributes. Reference images are from CelebA [1]. The number is the KL Divergence between <strong><em>Distribution of Generated Images vs. Ideal Uniform Distribution<\/em><\/strong>.<\/figcaption><\/figure>\n\n\n\n<p>We evaluate 5 text prompts in Figure 1\u2014 \u201ca headshot of a {person, professor, doctor, worker, firefighter}\u201d \u2014 and sample 200 images per prompt for each attribute, resulting in 40K generated images. We highlight the averaged results across 5 prompts of 6 attributes. ITI-GEN achieves near-perfect performance on balancing each binary attribute and multiple attributes, justifying our motivation: using separate inclusive tokens is beneficial in generating images that are uniformly distributed across attribute categories. <\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"323\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.09.18-AM-1024x323.png\" alt=\"\" class=\"wp-image-253\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.09.18-AM-1024x323.png 1024w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.09.18-AM-300x95.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.09.18-AM-768x242.png 768w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.09.18-AM.png 1066w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\"><em>Figure 2. <\/em>Qualitative results of the combination of four binary attributes (the last column in Figure 1). The input prompt (T ) is \u201ca headshot of a person\u201d. By using the learned inclusive tokens, ITI-GEN can inclusively generate images with all attribute combinations. Images across each tuple are sampled using the same random seed.<\/figcaption><\/figure>\n\n\n\n<p>In Figure 2, we show the qualitative results of our method in debiasing 4 binary attributes simultaneously. The results show that our method perform well in debiasing multiple attributes.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Multi-Category Attributes<\/strong><\/h2>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"740\" height=\"284\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.28.57-AM.png\" alt=\"\" class=\"wp-image-256\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.28.57-AM.png 740w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.28.57-AM-300x115.png 300w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\"><em>Figure 3. Multi-category distribution with \u201ca headshot of a person\u201d. The generated images from ITI-GEN are more uniformly distributed across different sub-groups than the baseline Stable Diffusion. See Figure 4 for qualitative results.<\/em><\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"740\" height=\"534\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.29.07-AM.png\" alt=\"\" class=\"wp-image-257\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.29.07-AM.png 740w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.29.07-AM-300x216.png 300w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\"><em>Figure 4. Results of ITI-GEN on multi-category attributes for Gender\u00d7Age (Figure 4(a)) and Gender\u00d7Skin Tone (Figure 4(b)). Examples are randomly picked with \u201ca headshot of a person\u201d.<\/em><\/figcaption><\/figure>\n\n\n\n<p>We further investigate multi-category attributes including perceived age and skin tone.<br>Specifically, we consider two challenging settings: (1) Perceived Gender \u00d7 Age (Figure 4(a)), and (2) Perceived Gender \u00d7 Skin Tone (Figure 4(b)). ITI-GEN achieves inclusiveness across all setups, especially on extremely under-represented categories for age (&lt; 10 and &gt; 50 years old in Figure 4(a)). More surprisingly (Figure 4(b)), ITI-GEN can leverage synthetic images (from FAIR) and jointly learn from different data sources (CelebA for gender and FAIR for skin tone), demonstrating great potential for bootstrapping inclusive data generation with graphics engines.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Other Domain<\/strong><\/h2>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"836\" height=\"808\" src=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.39.35-AM.png\" alt=\"\" class=\"wp-image-260\" srcset=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.39.35-AM.png 836w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.39.35-AM-300x290.png 300w, https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-08-at-12.39.35-AM-768x742.png 768w\" sizes=\"auto, (max-width: 706px) 89vw, (max-width: 767px) 82vw, 740px\" \/><figcaption class=\"wp-element-caption\"><br><em>Figure 5.<\/em> ITI-GEN with perception attributes (\u201cColorfulness\u201d) on scene images. The tokens of \u201ccolorfulness\u201d are trained with \u201ca photo of a natural scene\u201d and applied to \u201can alien pyramid landscape&#8230; \u201d in this example. ITI-GEN (bottom) enables the baseline Stable Diffusion (top) to generate images with different levels of colorfulness.<br><\/figcaption><\/figure>\n\n\n\n<p>Besides human faces, we apply ITI-GEN to another domain: scene images. We claim that the inclusive text-to-image generation accounts for attributes from not only humans but also scenes, objects, or even environmental factors. Specifically, we use images from LHQ [2] as guidance to learn inclusive tokens and generate images with diverse subjective perception attributes. As illustrated in Figure 5, ITI-GEN can enrich the generated images to multiple levels of colorfulness, justifying the generalizability of our method to the attributes in different domains.<\/p>\n\n\n\n<p>In conclusion, the results above demonstrate the effectiveness of our method in debiasing different attributes and different scenarios.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>References<\/strong>:<\/h2>\n\n\n\n<p>[1].\u00a0Ziwei Liu et al. &#8220;Deep learning face attributes in the wild.&#8221;\u00a0In ICCV, 2015<\/p>\n\n\n\n<p>[2]. Ivan Skorokhodov, et al. &#8220;Aligning latent and image spaces to connect the unconnectable.&#8221; In ICCV, 2021.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>In this section, we validate our method (ITI-Gen) for inclusive text-to-image generation on various attributes and scenarios. Single Binary &amp; Multiple Attributes To demonstrate the capability of our method to sample images with a variety of face attributes, we construct 40 distinct reference image sets based on attributes from CelebA [1]. Also, given multiple reference &hellip; <\/p>\n<p class=\"link-more\"><a href=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;Experiments&#8221;<\/span><\/a><\/p>\n","protected":false},"author":147,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-42","page","type-page","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Experiments - FairToken: Learning Fair Text Representations<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Experiments - FairToken: Learning Fair Text Representations\" \/>\n<meta property=\"og:description\" content=\"In this section, we validate our method (ITI-Gen) for inclusive text-to-image generation on various attributes and scenarios. Single Binary &amp; Multiple Attributes To demonstrate the capability of our method to sample images with a variety of face attributes, we construct 40 distinct reference image sets based on attributes from CelebA [1]. Also, given multiple reference &hellip; Continue reading &quot;Experiments&quot;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/\" \/>\n<meta property=\"og:site_name\" content=\"FairToken: Learning Fair Text Representations\" \/>\n<meta property=\"article:modified_time\" content=\"2023-05-08T04:45:06+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"4 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/introduction\\\/\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/introduction\\\/\",\"name\":\"Experiments - FairToken: Learning Fair Text Representations\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/introduction\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/introduction\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/wp-content\\\/uploads\\\/sites\\\/73\\\/2023\\\/05\\\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png\",\"datePublished\":\"2022-12-20T00:17:52+00:00\",\"dateModified\":\"2023-05-08T04:45:06+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/introduction\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/introduction\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/introduction\\\/#primaryimage\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/wp-content\\\/uploads\\\/sites\\\/73\\\/2023\\\/05\\\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png\",\"contentUrl\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/wp-content\\\/uploads\\\/sites\\\/73\\\/2023\\\/05\\\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/introduction\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Experiments\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/#website\",\"url\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/\",\"name\":\"FairToken: Learning Fair Text Representations\",\"description\":\"Unbiased Text-to-Image Generative Models | Students: Siqi Chai, Xuanbai Chen | Advisor: Fernando De la Torre\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/mscvprojects.ri.cmu.edu\\\/2023team2\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Experiments - FairToken: Learning Fair Text Representations","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/","og_locale":"en_US","og_type":"article","og_title":"Experiments - FairToken: Learning Fair Text Representations","og_description":"In this section, we validate our method (ITI-Gen) for inclusive text-to-image generation on various attributes and scenarios. Single Binary &amp; Multiple Attributes To demonstrate the capability of our method to sample images with a variety of face attributes, we construct 40 distinct reference image sets based on attributes from CelebA [1]. Also, given multiple reference &hellip; Continue reading \"Experiments\"","og_url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/","og_site_name":"FairToken: Learning Fair Text Representations","article_modified_time":"2023-05-08T04:45:06+00:00","og_image":[{"url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png","type":"","width":"","height":""}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/","url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/","name":"Experiments - FairToken: Learning Fair Text Representations","isPartOf":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/#website"},"primaryImageOfPage":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/#primaryimage"},"image":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/#primaryimage"},"thumbnailUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png","datePublished":"2022-12-20T00:17:52+00:00","dateModified":"2023-05-08T04:45:06+00:00","breadcrumb":{"@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/#primaryimage","url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png","contentUrl":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-content\/uploads\/sites\/73\/2023\/05\/Screen-Shot-2023-05-07-at-11.52.05-PM-1024x454.png"},{"@type":"BreadcrumbList","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/introduction\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/"},{"@type":"ListItem","position":2,"name":"Experiments"}]},{"@type":"WebSite","@id":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/#website","url":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/","name":"FairToken: Learning Fair Text Representations","description":"Unbiased Text-to-Image Generative Models | Students: Siqi Chai, Xuanbai Chen | Advisor: Fernando De la Torre","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"}]}},"_links":{"self":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-json\/wp\/v2\/pages\/42","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-json\/wp\/v2\/users\/147"}],"replies":[{"embeddable":true,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-json\/wp\/v2\/comments?post=42"}],"version-history":[{"count":32,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-json\/wp\/v2\/pages\/42\/revisions"}],"predecessor-version":[{"id":263,"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-json\/wp\/v2\/pages\/42\/revisions\/263"}],"wp:attachment":[{"href":"https:\/\/mscvprojects.ri.cmu.edu\/2023team2\/wp-json\/wp\/v2\/media?parent=42"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}