{"id":245046,"date":"2024-07-18T17:14:07","date_gmt":"2024-07-18T17:14:07","guid":{"rendered":"https:\/\/michigandigitalnews.com\/index.php\/2024\/07\/18\/apple-denies-using-youtube-content-to-train-apple-intelligence\/"},"modified":"2025-06-25T17:14:28","modified_gmt":"2025-06-25T17:14:28","slug":"apple-denies-using-youtube-content-to-train-apple-intelligence","status":"publish","type":"post","link":"https:\/\/michigandigitalnews.com\/index.php\/2024\/07\/18\/apple-denies-using-youtube-content-to-train-apple-intelligence\/","title":{"rendered":"Apple denies using YouTube content to train Apple Intelligence"},"content":{"rendered":"<p> [ad_1]<br \/>\n<br \/><img decoding=\"async\" src=\"https:\/\/readwrite.com\/wp-content\/uploads\/2024\/07\/a-stunning-cinematic-visual-of-an-apple-logo-trans-yBrGVavS3i6SkyxauFezw-xhPjRnu3Tkex1Ehg_9byHw-900x506.jpeg\" \/><\/p>\n<div>\n<p>Apple has denied using an unethically collected dataset from EleutherAI to train its flagship artificial intelligence (AI) product, Apple Intelligence. However, they state they have used the dataset for another AI model.<\/p>\n<p>After it was revealed this week that a company called EleutherAI used a dataset containing hundreds of thousands of YouTube video captions to create a dataset to aid in AI training, Apple <a href=\"https:\/\/appleinsider.com\/articles\/24\/07\/18\/apple-intelligence-wasnt-trained-on-stolen-youtube-videos\" target=\"_blank\" rel=\"noopener\">spoke to Apple Insider<\/a>, denying that EleutherAI\u2019s \u2018Pile\u2019 was used to train <a href=\"https:\/\/readwrite.com\/apple-intelligence-everything-you-need-to-know-about-your-iphones-new-big-brain\/\">Apple Intelligence<\/a>.<\/p>\n<p>However, they confirmed that \u2018the Pile\u2019 was used when developing the open-source OpenELM models released earlier this year.<\/p>\n<h2>What is EleutherAI\u2019s \u2018the Pile\u2019?<\/h2>\n<p>EleutherAI is a non-profit organization that wants to make AI research and development more accessible to companies outside of the huge tech firms we see primarily working on huge AI models like OpenAI.<\/p>\n<p>One of the ways they do this is by providing training datasets for large language models and other AI applications. However, instead of <a href=\"https:\/\/readwrite.com\/openai-seeks-media-licensing-for-language-models\/\">paying licensing fees to access data<\/a>, or entering into <a href=\"https:\/\/readwrite.com\/openai-and-news-corp-sign-major-deal-to-boost-chatgpt\/\" target=\"_blank\" rel=\"noopener\">partnerships to use data from sources<\/a>, EleutherAI scrapes the web to obtain its data. This includes the captions from over 170,000 YouTube videos.<\/p>\n<p>\u2018The Pile\u2019 is the result of this \u2013 a huge corpus of unethically sourced training data is intended to lower the barrier to entry for smaller firms to enter the AI market. However, larger companies have also made use of the dataset.<\/p>\n<h2>What is Apple\u2019s OpenELM?<\/h2>\n<p>Although they did not use \u2018the Pile\u2019 to train Apple Intelligence (and claim Apple Intelligence models were trained \u201con licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web crawler,\u201d) Apple has admitted to using it to develop their OpenELM models.<\/p>\n<p>Apple released OpenELM in April. It was created for research purposes and is not used to power any of Apple Intelligence\u2019s functions or features. Apple has <a href=\"https:\/\/9to5mac.com\/2024\/07\/17\/apple-intelligence-openelm-training-youtube\/\" target=\"_blank\" rel=\"noopener\">told 9to5Mac<\/a> that they have no plans to expand on OpenELM or release any further versions of the tool.<\/p>\n<p><strong><em>Featured image credit: Apple<\/em><\/strong><\/p>\n<\/p><\/div>\n<p>[ad_2]<br \/>\n<br \/><a href=\"https:\/\/readwrite.com\/apple-denies-using-youtube-content-to-train-apple-intelligence\/\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>[ad_1] Apple has denied using an unethically collected dataset from EleutherAI to train its flagship artificial intelligence (AI) product, Apple Intelligence. However, they state they<\/p>\n","protected":false},"author":1,"featured_media":245047,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_monsterinsights_skip_tracking":false,"_uf_show_specific_survey":0,"_uf_disable_surveys":false,"footnotes":""},"categories":[152],"tags":[],"_links":{"self":[{"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/posts\/245046"}],"collection":[{"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/comments?post=245046"}],"version-history":[{"count":0,"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/posts\/245046\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/media\/245047"}],"wp:attachment":[{"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/media?parent=245046"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/categories?post=245046"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/michigandigitalnews.com\/index.php\/wp-json\/wp\/v2\/tags?post=245046"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}