<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Data Science Weekly Newsletter]]></title><description><![CDATA[In-depth look at the Data Science / Machine Learning / AI / Data Engineering world.]]></description><link>https://datascienceweekly.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!I8ji!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3294ef46-cb03-42ea-b7d6-9f8e8b0f41f6_253x253.png</url><title>Data Science Weekly Newsletter</title><link>https://datascienceweekly.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 01 Aug 2026 07:22:10 GMT</lastBuildDate><atom:link href="https://datascienceweekly.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[datascienceweekly.org, a service of DATAYOU, LLC]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[datascienceweekly@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[datascienceweekly@substack.com]]></itunes:email><itunes:name><![CDATA[Data Science Weekly]]></itunes:name></itunes:owner><itunes:author><![CDATA[Data Science Weekly]]></itunes:author><googleplay:owner><![CDATA[datascienceweekly@substack.com]]></googleplay:owner><googleplay:email><![CDATA[datascienceweekly@substack.com]]></googleplay:email><googleplay:author><![CDATA[Data Science Weekly]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Data Science Weekly - Issue 662]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-662</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-662</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 30 Jul 2026 23:18:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9eIT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #662<br>July 23, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://www.hughparry.com/blog/music-mesh/">How to make a mesh dance in real time to music</a></strong><br>Audio transforms into a frequency space, where filtering becomes easy. The nice surprise is that meshes have a frequency space too&#8230;Take a drum head. Hit it and it doesn&#8217;t wobble arbitrarily; it rings in a discrete set of standing waves whose shapes are determined entirely by the drum&#8217;s geometry. Mark Kac asked in 1966 whether you could run that backwards: <em>can one hear the shape of a drum?&#8230;</em>On a triangle mesh those standing waves are the eigenvectors of the discrete Laplacian, its manifold harmonics. Low eigenvalues are slow global wobbles, high eigenvalues fine ripples: the same low-to-high ordering an FFT gives you, over a surface instead of over time&#8230;With this you can let the song drive the shape and get Kac&#8217;s question the other way. instead of hearing a shape, you see a sound&#8230;<br></p></li></ul><ul><li><p><strong><a href="https://www.sharonlohr.com/blog/2026/7/20/random-selection-and-democracy">Random Selection and Democracy in Athens and Asimov</a></strong><br>Let&#8217;s look at some aspects of random selection as it relates to democracy through these examples. This post contains discussion questions and activities that could be used for a high school or undergraduate class discussion about the &#8220;fairness&#8221; of simple random, stratified, and systematic sampling &#8212; and of nonrandom selection methods. Learning objectives include deepening understanding of these sampling methods, distinguishing between proportional and disproportional allocation, calculating sampling weights for a stratified sample, and relating advantages and disadvantages of methods for political candidate selection to statistical principles of randomization. The footnotes contain some questions for an advanced undergraduate or graduate-level sampling class&#8230;</p><p></p></li><li><p><strong><a href="https://equidistance.io/londons-most-equidistant-pub/">London&#8217;s most equidistant pub</a></strong><br>London&#8217;s most equidistant pub is <em>The Greene Man</em> at 383 Euston Road, on the corner of Great Portland Street. Starting from where each of the 33 boroughs&#8217; residents actually live, and arriving by 7pm on an ordinary September Thursday by public transport, journeys to it span 47 minutes between the nearest borough and the furthest, and no borough needs more than 70. No pub in London does better on both measures&#8230;.The bigger finding is what nobody manages: not one of London&#8217;s 3,170 pubs is within an hour of all 33 boroughs&#8230;.</p></li></ul><p>.</p><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:892786}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9eIT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9eIT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png 424w, https://substackcdn.com/image/fetch/$s_!9eIT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png 848w, https://substackcdn.com/image/fetch/$s_!9eIT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png 1272w, https://substackcdn.com/image/fetch/$s_!9eIT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9eIT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png" width="601" height="339.37587412587413" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6c514fc-e416-4204-bc66-b6012b286748_1144x646.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:646,&quot;width&quot;:1144,&quot;resizeWidth&quot;:601,&quot;bytes&quot;:71928,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/209185353?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9eIT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png 424w, https://substackcdn.com/image/fetch/$s_!9eIT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png 848w, https://substackcdn.com/image/fetch/$s_!9eIT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png 1272w, https://substackcdn.com/image/fetch/$s_!9eIT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6c514fc-e416-4204-bc66-b6012b286748_1144x646.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://blog.doubleword.ai/you-could-have-come-up-with-kimi-delta-attention">You Could Have Come Up With Kimi Delta Attention</a></strong><br>In this post we are going to walk through the DeltaNet family of linear attention variants, two of which are used by the latest Qwen and Kimi model families, and show how you might have arrived at the same equations by asserting simple things about your hidden state. That is the route we will take: softmax attention &#8194;&#8594;&#8194; linear attention &#8194;&#8594;&#8194; DeltaNet &#8194;&#8594;&#8194; Gated DeltaNet &#8194;&#8594;&#8194; KDA Only after deriving KDA will we turn to the recurrent and chunkwise Triton programs that execute it&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/analytics/comments/1v7uvh8/anyone_else_measure_team_health_by_how_long_it/">Anyone else measure team health by how long it takes a non-technical person to answer a basic data question? [Reddit]</a></strong></p><p>had a client once where the ops manager could just... answer things. revenue by region, churn this month, whatever. no ticket, no slack message to the data team, just opened metabase and pulled it herself i&#8217;ve been thinking about that a lot lately because most places i go into are nothing like that. analyst gets pinged for stuff that really shouldn&#8217;t need an analyst&#8230;</p><p></p></li><li><p><strong><a href="https://emschwartz.me/binary-vector-embeddings-are-so-cool/">Binary vector embeddings are so cool</a><br></strong>Vector embeddings by themselves are pretty neat. Binary quantized vector embeddings are extra impressive. In short, they can retain 95+% retrieval accuracy with 32x compression and ~25x retrieval speedup. Let&#8217;s get into how this works and why it&#8217;s so crazy&#8230;<br></p></li><li><p><strong><a href="https://www.statforbiology.com/posts/nls_usefulequations">Some useful equations for biological processes</a><br></strong>Because mathematical modeling plays such a central role in biology and many other scientific disciplines, biologists need to be familiar with the most important mathematical functions. More importantly, they need to be able to &#8220;read&#8221; these functions and use their parameters to describe, interpret, and quantify biological processes. With this aim in mind, I have compiled a collection of the mathematical functions most commonly encountered in biology, explaining the meaning of their parameters, with particular emphasis on their biological interpretation rather than their mathematical properties&#8230;<br></p></li><li><p><strong><a href="https://www.statforbiology.com/posts/nls_modelfitting">A collection of self-starters for nonlinear regression in R</a></strong></p><p>Usually, the first step in every nonlinear regression analysis is to select the function that best describes the phenomenon under study. The next step is to fit this function to the observed data, possibly by using some sort of nonlinear least squares algorithm&#8230;for the second step, the main problem is that nonlinear least squares algorithms are iterative, in the sense that they start from some initial guesses for the model parameters, which are continuously improved until the least squares solution is approximately reached. Quite often, providing such initial guesses for all model parameters becomes a problem: if our guesses are not close enough to the least squares estimates, the algorithm may stall and fail to converge. Or, even worse, it may converge to the wrong solution. How do we obtain good initial guesses for the model parameters? This is not easily accomplished, especially for students and practitioners. This is where self-starters come in handy&#8230;<br></p></li><li><p><strong><a href="https://evanhahn.com/prefer-strict-tables-in-sqlite/?utm_source=hackernewsletter&amp;utm_medium=email&amp;utm_term=data">Prefer STRICT tables in SQLite</a></strong><br>In short: I prefer strict tables in SQLite because they avoid some datatype problems, such as putting text in number columns. SQLite has a feature that I think is underrated: strict tables. Strict tables help enforce rigid typing, preventing mistakes like putting text into integer columns. I like them, and wrote this post to promote their use!&#8230;<code><br></code></p></li><li><p><strong><a href="https://www.reddit.com/r/statistics/comments/1v8n86l/question_why_did_data_scientists_choose_rows_to/">Why did data scientists choose Rows to be observations? [Reddit]</a><br></strong>I&#8217;m a math guy so i don&#8217;t know sh*t about stats but in linalg we&#8217;re learning about covariance matrices, PCA, and SVD. In math we prefer columns to be observations because&#8230; well tradition, notational easy (Ax=b instead of xA=b if we used row vectors) and cuz we care abt linear transformations etc and its easier to think about columns. Idk. Mostly tradition tho. But then why did data scientists break the mold? What benefit did row observations possibly serve that Col vectors/observations couldn&#8217;t give?? I am jumping back and forth between notation and conventions and it&#8217;s hard to keep up&#8230;<br></p></li><li><p><strong><a href="https://posh.wiki/blog/2026-07-20-hoop-sizing/">Cross-stitch prep for math nerds</a><br></strong>I&#8217;ve done quite a bit of fibre arts over the years, and surprisingly enough, picking up the right supplies in the right quantities is probably the hardest part. This post details how to figure out exactly what supplies you need via mathematics and geometry&#8230;<br></p></li><li><p><strong><a href="https://lcamtuf.substack.com/p/tech-note-making-your-own-v-i-plots">Tech note: making your own V-I plots at home&#8230;Or, how to lose money with friends</a></strong></p><p>When working on my latest book, The Secret Life of Circuits, I wanted to keep the artwork real. My beef with the diagrams in popular electronics textbooks and online tutorials is that most of them are fake. At best, they&#8217;re retraced from ancient texts; at worst, they&#8217;re sketched from memory and can be charitably described as &#8220;inspired by true events&#8221;...<br></p></li><li><p><strong><a href="https://softwaredoug.com/blog/2026/07/29/just-brute-force-embeddings">Just brute force your embeddings</a></strong><br>I work with a lot of teams that don&#8217;t need the complexity of a vector database. They have ~1m documents to search. They have low query traffic, and write their embeddings all up front. They don&#8217;t need to buy a multi-million dollar vector database, or spend 6 months learning to operate it. For low enough n, just brute force the embeddings until you can&#8217;t bear to. As fellow search traveler Jo Kristian Bergum says &#8220;an exhaustive search may be all you need&#8221;&#8230;<br></p></li><li><p><strong><a href="https://ericmjl.github.io/blog/2026/7/23/going-bayesian-automates-data-analysis/">Going Bayesian automates your manual data analysis</a></strong></p><p>People usually talk about Bayesian modeling in terms of richer uncertainty estimates, prior knowledge, and posterior distributions. All true, and all worth the switch on their own. But there&#8217;s a benefit almost nobody talks about. Going Bayesian automates away the manual data analysis you&#8217;d otherwise do by hand. And once you experience it, you wonder how you ever worked any other way. Let me show you what I mean&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1v5z7ue/how_do_you_decide_whether_a_data_science_problem/">How do you decide whether a data science problem really needs machine learning? [Reddit]</a><br></strong>In your experience, what factors help you decide between using a simple analytical approach and building a machine learning model? I&#8217;d love to hear the reasoning behind your decision-making process&#8230;<br></p></li><li><p><strong><a href="https://www.staszewski.xyz/blog/recursive-filters/">Recursive Filters: SMA, EMA, Low&#8209;Pass, and a Tiny Kalman</a><br></strong>While exploring optimizers I fell down the rabbit hole of recursive filters. This post is a compact, practical tour of smoothers you can use when measurements are noisy but latency and compute are tight. We&#8217;ll keep the math minimal, the intuition high, and focus on when and why to use SMA, EMA/low&#8209;pass, and a tiny 1D Kalman&#8230;</p></li></ul><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1v4l44b/what_do_todays_data_science_graduates_commonly/">What Do Today&#8217;s Data Science Graduates Commonly Lack? [Reddit]</a></strong></p></li><li><p><strong><a href="https://suzyahyah.github.io/machine%20learning/2026/06/27/trouble-with-time-series.html">The Unreasonable Difficulty of Time Series Forecasting</a></strong></p></li><li><p><strong><a href="https://www.empirical.health/blog/non-invasive-glucose-monitoring-wearables/">Why non-invasive glucose monitoring is hard</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** You can find last week's issue #661 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-661">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://www.fharrell.com/post/cpm/">The Unifying Capabilities of Cumulative Probability Semiparametric Models</a></strong></p></li><li><p><strong><a href="https://quantixed.org/2026/07/21/colorblind-ii-microscopy-images-and-colour-blindness/">Colorblind II: Microscopy images and colour blindness</a></strong></p></li><li><p><strong><a href="https://jvns.ca/blog/2026/07/17/learning-about-running-sqlite/">Learning a few things about running SQLite</a></strong></p></li><li><p><strong><a href="https://berkeleyautomation.github.io/robovista/">RoboVista - Evaluating Vision-Language Models for Diverse Robot Applications</a></strong></p></li></ul><p>.</p><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 661]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-661</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-661</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 23 Jul 2026 21:57:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oCaL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #661<br>July 23, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://www.empirical.health/blog/non-invasive-glucose-monitoring-wearables/">Why non-invasive glucose monitoring is hard</a></strong><br>Continuous glucose monitoring on Apple Watch and other smartwatches has been &#8220;5-7 years away&#8221; for roughly a decade&#8230;I trained one of the first deep neural networks to detect signs of diabetes with consumer health sensors. This post will explain what makes glucose sensing so hard, what hardware and machine learning techniques have been tried, and try to describe which research techniques are actually feasible on a consumer device like an Apple Watch, Pixel, Oura, or Samsung Watch. Let&#8217;s start by explaining why an already-launched feature, blood oxygen sensing, actually works in practice using relatively cheap optical sensors&#8230;<br></p></li></ul><ul><li><p><strong><a href="https://jakubnowosad.com/posts/2026-09-08-erdkunde/">Navigating Challenges in Spatial Machine Learning</a></strong><br>Spatial machine learning has become a standard tool for producing environmental and geographic prediction maps. It is now relatively (technically) easy to combine field observations with remote sensing, climate, terrain, or other predictor layers and fit a strong machine learning model. The harder question is whether the resulting map is reliable, transferable, and reproducible&#8230;Spatial dependence, clustered and biased sampling, heterogeneous landscapes, and domain transfer all affect how models should be evaluated and interpreted. A model can appear accurate under a standard validation approach and still be unreliable where predictions are needed&#8230;</p><p></p></li><li><p><strong><a href="https://cbowdon.github.io/posts/gymflation/">Exploring Gymflation with AI</a></strong><br>You might be familiar with super hero inflation? Over time, super hero physiques on screen have become increasingly exaggerated. Batman&#8217;s progression from Adam West to Ben Affleck is a great example of this&#8230;Gymflation is much the same idea. As gym culture has skyrocketed, it seems like so too have people&#8217;s feats of strength - as recorded on social media. Going on YouTube or Instagram one gets the feeling that a 200kg deadlift is really rather average nowadays. But is it really true? Or are we just feeling the effects of the Algorithm, pushing extreme examples in our blue-lit faces? What follows is a little project to find out&#8230;</p></li></ul><p>.</p><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:846549}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oCaL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oCaL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png 424w, https://substackcdn.com/image/fetch/$s_!oCaL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png 848w, https://substackcdn.com/image/fetch/$s_!oCaL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png 1272w, https://substackcdn.com/image/fetch/$s_!oCaL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oCaL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png" width="601" height="369.84615384615387" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84701944-726a-44fa-9d2d-e5870994e099_1144x704.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:704,&quot;width&quot;:1144,&quot;resizeWidth&quot;:601,&quot;bytes&quot;:73971,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/208256027?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oCaL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png 424w, https://substackcdn.com/image/fetch/$s_!oCaL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png 848w, https://substackcdn.com/image/fetch/$s_!oCaL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png 1272w, https://substackcdn.com/image/fetch/$s_!oCaL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84701944-726a-44fa-9d2d-e5870994e099_1144x704.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://www.nber.org/papers/w34952">Easy A&#8217;s, Less Pay: The Long-Term Effects of Grade Inflation</a></strong><br>Average grades continue to rise in the United States, raising the question of how grade inflation impacts students. We provide comprehensive evidence on how teacher grading practices affect students' long-run success. Using administrative high school data from Los Angeles and from Maryland that is linked to postsecondary and earnings records, we develop and validate two teacher-level measures of grade inflation: one measuring average grade inflation and another measuring a teacher's propensity to give a passing grade&#8230;The cumulative impact is economically significant: a teacher with one standard deviation higher average grade inflation reduces the present discounted value of lifetime earnings of their students by $213,872 per year&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1v4l44b/what_do_todays_data_science_graduates_commonly/">What Do Today&#8217;s Data Science Graduates Commonly Lack? [Reddit]</a></strong></p><p>I often read comments from hiring managers and interviewers saying they&#8217;re disappointed with recent data science graduates. I&#8217;m curious, what do you think these graduates are lacking? If someone wants to become a data scientist, what skills should they focus on? Strong software engineering skills? Math and statistics? Something else?&#8230;</p><p></p></li><li><p><strong><a href="https://www.statsignificant.com/p/how-hollywood-stopped-making-movies">How Hollywood Stopped Making Movies in Hollywood</a><br></strong>The economics behind where movies are filmed, why Hollywood left Los Angeles, and whether audiences can tell the difference&#8230;<br></p></li><li><p><strong><a href="https://suzyahyah.github.io/machine%20learning/2026/06/27/trouble-with-time-series.html">The Unreasonable Difficulty of Time Series Forecasting</a><br></strong>I&#8217;ve been thinking recently about what makes time series forecasting problems so difficult compared to other sequence learning tasks or IID Machine Learning problems&#8230;for many ML problems, we don&#8217;t have full information about the casual factors either yet are able to do something reasonable, or at least better than a random walk. Hence, this post investigates the nature of the forecasting problem and what makes it so much more difficult than classical machine learning&#8230;<br></p></li><li><p><strong><a href="https://thegustafson.com/series">Holding the LLM Stack in Your Head</a></strong></p><p><span>A dependency-ordered walk through the modern LLM stack, from the linear algebra under a single attention head, through training and inference, out to agent protocols shipping in 2026. Ten arcs, eighty-odd posts. The goal isn't rigor, it's intuition that survives contact with real systems&#8230;</span><br></p></li><li><p><strong><a href="https://golfcoursewiki.substack.com/p/how-to-bake-a-dispersion-pattern">How to Bake a [Golf Shot] Dispersion Pattern from Scratch</a></strong><br>I&#8217;ve built a realistic golf shot dispersion simulator and I&#8217;ve shared it for anyone to use&#8230;I wanted to write about variance in golf here, but when I was reviewing my theoretical, idealized dispersion patterns, I realized it was effectively pointless to start discussing variance in golf without generating realistic dispersion patterns. These golf shot dispersion patterns represent the inherent variance in the game, so we need the model to reflect them before looking at birdies and blow-ups&#8230;<code><br></code></p></li><li><p><strong><a href="https://quantixed.org/2022/05/19/colorblind-checking-figure-accessibility-for-colour-blind-people/">Colorblind: Checking figure accessibility for colour blind people</a><br></strong>When preparing images for publication, it is good practice to check how accessible they are for colour blind people. Using a simple bit of code, it is possible to check an image &#8211; or a whole figure &#8211; in ImageJ for accessibility. For example, Figure 1 from our recent paper. Originally looked like this:..<br></p></li><li><p><strong><a href="https://blog.google/innovation-and-ai/technology/research/understanding-the-ai-economy/">Understanding the AI economy</a><br></strong>Google&#8217;s ATLAS is an expansive look at how people are using AI at work and in day-to-day life&#8230;Google is launching the first iteration of the AI &amp; Economy ATLAS (Activity, Task, Landscape, and Adoption Study), an ongoing, large-scale, de-identified study of how people are using Google&#8217;s AI products and tools. ATLAS&#8217;s first dataset (v1.0) is built from 15 million aggregated and de-identified human-AI interactions across the Gemini App, AI Mode, and the Gemini API, which together are used by more than 1 billion people monthly. ATLAS v1.0 insights span more than 150 countries, 140 languages, 800 occupations, and 4,000 tasks; ATLAS is the most comprehensive look to date at how real people are using AI at scale&#8230;<br></p></li><li><p><strong><a href="https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2023.1219945/full">Handwriting but not typewriting leads to widespread brain connectivity: a high-density EEG study with implications for the classroom</a></strong></p><p>Our findings suggest that the spatiotemporal pattern from visual and proprioceptive information obtained through the precisely controlled hand movements when using a pen, contribute extensively to the brain&#8217;s connectivity patterns that promote learning&#8230;<br></p></li><li><p><strong><a href="https://sinja.io/blog/data-landscape-guide-for-developers">Guide to data tools landscape for developers</a></strong><br>Found yourself on a data project and have no idea what they all are talking about? Feel excluded from all the fun discussions in the office kitchen? If only there were a humongous guide going over all the concepts and buzzwords.&#8230;.In this article, we will briefly go over the data lifecycle: where data comes from, how it&#8217;s handled, how it&#8217;s stored, and how it&#8217;s displayed. You will understand to which stage each particular tool belongs and which tasks it solves for people working with data. We won&#8217;t go into details about setting up each tool or comparing tools inside the same class in-depth. Trust me, the article will be pretty lengthy as is, even without going into any of those details&#8230;<br></p></li><li><p><strong><a href="https://marclamberts.medium.com/scouting-uncertainty-adding-value-to-data-scouting-results-f68d8c148385">Scouting Uncertainty: adding value to data scouting results</a></strong></p><p>I have been looking at spreadsheets for football scouting for over a decade, and every now and then, my perception completely changes&#8230;Lately, I&#8217;ve been creating more and more on a meta level, mostly in data engineering. In that light, I wanted to share something I have introduced in my day-to-day data scouting when working with aggregated data: scouting uncertainty. The idea is to add an uncertainty score to the scouting score to add a level of trust towards the data. If the data is trustworthy and uncertainty is low, the quality of data scouting for that specific player will be higher&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/dataengineering/comments/1utih2r/ai_as_an_etl_and_report_builder_im_tired/">AI as an ETL and Report Builder? I&#8217;m tired. [Reddit]</a><br></strong>We have been developing a Data Platform (IaC, CI/CD, orchestration, data quality, governance, the works). Everything is already set-up except for the business logic. Quite understandable since we built everything from FOSS about 2 months ago and I&#8217;m the only data platform engineer/data engineer in the company. They aren&#8217;t also keen on spending money on managed solutions. Now, a director is pushing to scrap our project in favor of an AI as an ETL solution. Basically, use skills and AI to generate reports from source systems and have AI use python, pandas and SQL to generate reports. This AI as an ETL couldn&#8217;t get out of the demo phase because of data quality issues&#8230;.<br></p></li><li><p><strong><a href="https://datascienceconfidential.github.io//r/economics/book-reviews/2026/07/19/reading-radical-undertainty.html">Reading &#8216;Radical Uncertainty&#8217;</a><br></strong>I recently read John Kay and Mervyn King&#8217;s book Radical Uncertainty as part of an economics book club and I thought I would share some thoughts on the book here&#8230;Radical Uncertainty is about statistics, and especially about the problems with assuming that we can make probabilistic calculations while neglecting the possibility of things which we haven&#8217;t even contemplated happening. I was pretty fascinated by the first half of the book, but it does become quite long-winded in the second half, and repeats many of its points over and over again. The book makes a lot of good points, but I would also like to point out some mistakes which it makes when talking about statistical modelling&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://www.reddit.com/r/dataengineering/comments/1uvd60m/how_do_you_visualize_sql_in_your_head/">How do you visualize SQL in your head? [Reddit]</a></strong></p></li><li><p><strong><a href="https://www.theocharis.dev/blog/llm-critics-are-right-i-use-llms-anyway/">The LLM Critics Are Right. I Use LLMs Anyway.</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/dataengineering/comments/1uvss0i/if_you_had_to_rebuild_your_entire_data_platform/">If you had to rebuild your entire data platform today from scratch, what stack would you choose? [Reddit]</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** You can find last week's issue #660 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-660">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://postgres.saneengineer.com/">Size a PostgreSQL server on AWS</a></strong></p></li><li><p><strong><a href="https://artem.krylysov.com/blog/2026/07/23/how-mvcc-and-transactions-work-in-rocksdb/">How MVCC and Transactions Work in RocksDB</a></strong></p></li><li><p><strong><a href="https://www.bbc.co.uk/programmes/w3ct998z">How much luck do you need to win the World Cup? - More or Less Investigating the role of random chance in football</a></strong></p></li><li><p><strong><a href="https://jakubnowosad.com/posts/2026-07-21-user/">A world still to be mapped: reflections on geocomputation in R: takeaways from the talk and workshop at UseR! 2026</a></strong></p></li></ul><p>.</p><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 660]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-660</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-660</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Fri, 17 Jul 2026 00:42:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Pc-H!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #660<br>July 17, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://www.reddit.com/r/dataengineering/comments/1uvd60m/how_do_you_visualize_sql_in_your_head/">How do you visualize SQL in your head? [Reddit]</a></strong><br>I'm a software engineer that currently helping to build data team, so I work a lot as an "analyst" and build a bunch of data models&#8230;During code review, I&#8217;m hardly able to tell potential bugs (like joining using wrong key, potential row explosion, etc) at a glance, unlike when reviewing eg: python code. Atm, I&#8217;m 50:50 asking claude to generate me a simple viz that could help me trace what happens in each cte transformation. Not always useful, but it slightly helps me. How do you personally tackle this? Do you have an easier mental model that you want to share?&#8230;<br></p></li></ul><ul><li><p><strong><a href="https://wc26.bogachev.fr/index.html"><span>Football DataPortraits</span></a></strong><br>It's an impression, but one built entirely from data.<span> Nothing is staged: each match is reconstructed from roughly 1,500 recorded events &#8212; every touch, pass, shot and card&#8230;</span></p><p></p></li><li><p><strong><a href="https://codebynight.dev/posts/arima-is-boring-and-that-is-why-i-still-like-it/">ARIMA Is Boring, and That Is Why I Still Like It</a></strong><br>ARIMA is old, limited, and a little annoying. I still like it because it makes the assumptions and uncertainty in a forecast harder to ignore&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:802875}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Pc-H!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Pc-H!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png 424w, https://substackcdn.com/image/fetch/$s_!Pc-H!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png 848w, https://substackcdn.com/image/fetch/$s_!Pc-H!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png 1272w, https://substackcdn.com/image/fetch/$s_!Pc-H!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Pc-H!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png" width="600" height="321.54115586690017" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:612,&quot;width&quot;:1142,&quot;resizeWidth&quot;:600,&quot;bytes&quot;:65135,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/207360283?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Pc-H!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png 424w, https://substackcdn.com/image/fetch/$s_!Pc-H!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png 848w, https://substackcdn.com/image/fetch/$s_!Pc-H!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png 1272w, https://substackcdn.com/image/fetch/$s_!Pc-H!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74e4f8ac-d283-4f29-bd56-5164fbf937c4_1142x612.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://www.theocharis.dev/blog/llm-critics-are-right-i-use-llms-anyway/">The LLM Critics Are Right. I Use LLMs Anyway.</a></strong><br>This week I was at Local-First Conf in Berlin, and the dissonance was everywhere&#8230;I spoke at that conference myself, and when I later talked to some of the people, they described the feeling as pretty similar to mine, which is a relief, because I know I am not alone with this&#8230;So this article is me trying to describe it. I&#8217;ll start by going through all of the fair and valid concerns about using LLMs, the things that would get the big round of applause. Then I will explain what makes me still use LLMs. And I&#8217;ll finish up with some of the patterns I found, in the hope that by giving concrete examples, others can step in as well and describe their experiences, so we can all come together and get a better understanding of this dissonance&#8230;.<br></p></li><li><p><strong><a href="https://www.reddit.com/r/dataengineering/comments/1uvss0i/if_you_had_to_rebuild_your_entire_data_platform/">If you had to rebuild your entire data platform today from scratch, what stack would you choose? [Reddit]</a></strong></p><p>Mid-sized company</p><p>Cloud-native</p><p>Batch + streaming</p><p>SQL-heavy analytics</p><p>Some ML workloads&#8230;</p><p></p></li><li><p><strong><a href="https://blog.haskell.org/enterprise-haskell-at-h-e-b/">Enterprise Haskell at H-E-B</a><br></strong>When I joined H-E-B back in 2018, I was entering a world far removed from any of my previous roles in technology. H-E-B is a retail company first, and the largest privately held company in Texas. Our customers care about full shelves, not the systems behind them, and for decades those systems delivered: rock-solid COBOL mainframes doing their job reliably, day in and day out&#8230;<br></p></li><li><p><strong><a href="https://github.com/alxndrTL/little-book-rl/">The Little Book of Reinforcement Learning</a><br></strong>This book is a short introduction to Reinforcement Learning, from the basics to applied algorithms&#8230;<br></p></li><li><p><strong><a href="https://ropensci.org/blog/2026/07/14/15yo-eunseop-kim/">From Peer Review to Mentorship: My rOpenSci Story</a></strong></p><p><span>I first came to rOpenSci in 2022, though at the time I barely knew what it was. I was getting a statistical package of mine ready to submit to the Journal of Statistical Software, and that is how I was pointed toward </span><a href="https://ropensci.org/software-review/">rOpenSci review</a><span>: the journal directs authors to rOpenSci&#8217;s statistical software standards, so going through the review looked like a convenient step along the way. At the time, my focus was on polishing the software for the journal submission, not on rOpenSci itself&#8230;</span><br></p></li><li><p><strong><a href="https://www.arenaphysica.com/publications/smith-charts">What is a Smith Chart?</a></strong><br>An interactive introduction to the chart at the heart of RF engineering&#8230;<code><br></code></p></li><li><p><strong><a href="https://latent-thought-flows.vercel.app/">Latent Thought Flows with Text Compression</a><br></strong>For images, audio, video, and actions, modern generative modeling increasingly shares one recipe: compress the signal into continuous latent tokens, train a generator to map noise to those tokens, and decode them back into the original domain. Language has been the exception&#8230;Our core message is simple: text can be compressed into a short sequence of continuous latents that remains useful for generation. This lets language share the same latent-generative recipe as the other modalities, while a text decoder handles surface realization&#8230;<br></p></li><li><p><strong><a href="https://gregorygundersen.com/blog/2025/10/01/large-language-models/">A History of Large Language Models</a><br></strong>I trace an academic history of some of the core ideas behind large language models, such as distributed representations, transducers, attention, the transformer, and generative pre-training&#8230;<br></p></li><li><p><strong><a href="https://www.nature.com/articles/s41586-026-10780-5">A Bayesian framework for longitudinal EHR and genetic discovery</a></strong></p><p>Electronic health records (EHRs) provide rich longitudinal disease histories, but existing methods for analyzing these data typically treat diseases in isolation and rarely integrate germline genetics. Here we present ALADYNOULLI, a Bayesian generative framework that jointly models longitudinal EHR diagnoses, age and polygenic risk to recover latent time-varying disease signatures and patient-specific signature loadings&#8230;<br></p></li><li><p><strong><a href="https://statmodeling.stat.columbia.edu/2026/07/15/a-ranked-choice-election-in-maine-using-voting-data-to-understand-preferences/">A ranked-choice election in Maine, USA: Using voting data to understand preferences</a></strong><br><span>Maine uses ranked choice voting (RCV) in primaries and federal elections, so voters could rank up to six choices for Governor. Using the </span>instant runoff<span> algorithm, candidates were sequentially dropped based on who had the fewest first-choice votes, and ballots were reallocated to each voter&#8217;s next-ranked choice&#8230;</span>But analyses of the individual ballots cast in the primary reveal a surprising mathematical fact: though she was eliminated before them, Bellows would have defeated either Shah  or Jackson in one-on-one elections&#8230;This unintuitive fact is a generalization of a well-known feature of ranked choice voting elections: it does not satisfy the Condorcet winner criterion&#8230;.<br></p></li><li><p><strong><a href="https://oaklab.ai/posts/learning-from-experience-instead-of-curated-datasets">Learning from experience instead of curated datasets</a></strong></p><p>Learning from experience is different from learning from curated datasets. Algorithms that learn from curated datasets can assume that all data is useful for learning. This is usually guaranteed by humans, who collect, clean, and filter raw data so that it is ready to be consumed by a learning algorithm. Experience, on the other hand, does not always contain learnable associations&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1usmwn8/all_these_layoffs_have_made_me_question_my_job/">All these layoffs have made me question my job search [Reddit]</a><br></strong>Right now I&#8217;m at a company that hasn&#8217;t done layoffs since maybe the financial crisis. I know how fortunate that is. But if I switch jobs, I could make an extra $50K. So I keep asking myself: is that extra 50K worth the instability that comes with tech jobs right now? What if I join a company and get laid off within a year?&#8230;<br></p></li><li><p><strong><a href="https://minsukchang.com/blog/2026-07-15-human-in-context.md.html">Users are Humans in Some Context</a><br></strong>If we model a human as a conditional probability distribution: <em><strong>P(action | context)</strong>,</em></p><p style="text-align: justify;">then the essence of understanding human behavior&#8212;and designing systems for them&#8212;lies in unpacking what actually constitutes that context and how the action is selected&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://arxiv.org/abs/2206.07867">A visual introduction to information theory</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1upl6er/managing_dealing_with_junior_data_scientists/">Managing/ Dealing with Junior Data Scientists? [Reddit]</a></strong></p></li><li><p><strong><a href="https://www.ethanrosenthal.com/2016/07/20/lets-talk-or/">I&#8217;m all about ML, but let&#8217;s talk about OR</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** You can find last week's issue #659 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-659">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://www.bbc.co.uk/programmes/p0nxwkcw">Does playing tennis make you live 9.7 years longer?</a></strong></p></li><li><p><strong><a href="https://arxiv.org/abs/2607.11362">Boolean queries are all you need?</a></strong></p></li><li><p><strong><a href="https://loganthrashercollins.substack.com/p/100-scientific-papers-ive-read-in">100 scientific papers I&#8217;ve read in full over the past year</a></strong></p></li><li><p><strong><a href="https://link.springer.com/article/10.1007/s13752-026-00544-9">What Lives? A Meta-Analysis of Diverse Opinions on the Definition of Life</a></strong></p></li><li><p><strong><a href="https://data-syn.github.io/">Domain-Aware Scaling Laws Uncover Data Synergy</a></strong></p></li></ul><p>.</p><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 659]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-659</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-659</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Fri, 10 Jul 2026 04:04:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!RRtN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #659<br>July 09, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://arxiv.org/abs/2206.07867">A visual introduction to information theory</a></strong><br>We present a visual, intuition-driven guide to key concepts in information theory. We show how entropy, mutual information, and channel capacity follow from basic probability, and how they determine the shortest possible encoding of a data source and the maximum rate of reliable communication through a noisy channel. Our presentation assumes only a familiarity with basic probability theory&#8230;<br></p></li></ul><ul><li><p><strong><a href="https://otexts.com/weird/">That&#8217;s weird! Anomaly detection using R</a></strong><br>This book is about tools and techniques for finding and understanding anomalies. We will begin with some simple data sets containing only one variable, and build up slowly to much more complicated data. We will cover popular but inadvisable methods to identify anomalies (pointing out their shortcomings), as well as more reliable and recommended approaches&#8230;.</p><p></p></li><li><p><strong><a href="https://jakubnowosad.com/posts/2026-07-03-ml4eo/">Rethinking Validation for Spatial Machine Learning: Takeaways from the Talk</a></strong><br>A summary of key points from my keynote and workshop at the Machine Learning for Earth Observation conference in Exeter (2026-06-22 and 2026-06-23)&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:755860}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RRtN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RRtN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png 424w, https://substackcdn.com/image/fetch/$s_!RRtN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png 848w, https://substackcdn.com/image/fetch/$s_!RRtN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png 1272w, https://substackcdn.com/image/fetch/$s_!RRtN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RRtN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png" width="610" height="382.06405693950177" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:704,&quot;width&quot;:1124,&quot;resizeWidth&quot;:610,&quot;bytes&quot;:74818,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/206388443?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RRtN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png 424w, https://substackcdn.com/image/fetch/$s_!RRtN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png 848w, https://substackcdn.com/image/fetch/$s_!RRtN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png 1272w, https://substackcdn.com/image/fetch/$s_!RRtN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb027899c-ee93-49ad-947a-ae4b8584e0e6_1124x704.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase">Benchmarking Coding Agents on Databricks&#8217; Multi-Million Line Codebase</a></strong><br>This article shares the results and methodology of the internal coding benchmark we built at Databricks, which evaluates tools on actual coding tasks our engineers performed on the Databricks codebase. Tasks featured edits against a multi-million line codebase covering many popular languages (Python, Go, Typescript, Scala, etc.) and both tasks and solutions were carefully reviewed to ensure accuracy. This isn&#8217;t meant to be comprehensive, but the exercise surfaced insights that have already made our engineering team meaningfully more efficient with coding agents. Below, you can see how models and harnesses scored on the overall benchmark&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/analytics/comments/1uo9we2/how_much_do_you_actually_trust_ai_output_for_real/">How much do you actually trust AI output for real reporting? [Reddit]</a></strong></p><p>Feels like everyone&#8217;s using ChatGPT/Copilot for something now, but I can&#8217;t tell whether it&#8217;s mostly &#8220;write me a first draft I&#8217;ll review&#8221; or something people are comfortable using for production reporting after proper validation. Anyone run into cases where AI-generated code looked completely reasonable but was subtly wrong? How much manual verification do you guys still do vs. just running with it?&#8230;</p><p></p></li><li><p><strong><a href="https://www.modular.com/blog/matrix-multiplication-on-nvidias-blackwell-part-1-introduction">Matrix Multiplication on Blackwell: Part 1 - Introduction</a><br></strong>In Part 1 (this blog post) we cover what a Matrix Multiplication (matmul) is, its importance for LLMs, and why we need to optimize it. Then we explain what a GPU is, GPU history since Ampere, and finally how to write a simple (not super performant) implementation of matmul on a GPU in 4 lines of Mojo. In part 2, we&#8217;ll explain the hardware instructions introduced in Blackwell GPUs, and continue improve on our kernels&#8217; performance to make it leverage the new hardware instructions. As we continue through the blog series, we will incrementally leverage new Blackwell features to improve our matmul implementation until the end of the series where we achieve performance that surpasses that of NVIDIA&#8217;s cuBLAS library&#8230;<br></p></li><li><p><strong><a href="https://hal.science/hal-05684645">Optimal Transport for Actuarial Science</a><br></strong>These lecture notes introduce optimal transport as a mathematical language for actuarial science. They treat losses, premiums, scores, reserves, capital scenarios, climate losses and lifetime distributions as probability measures that can be compared, transported, averaged, stressed and interpolated. The first part develops the main tools: couplings, push-forwards, discrete and continuous Kantorovich problems, duality, Wasserstein distances, quantile transport, barycenters, entropic regularization and statistical optimal transport. The second part applies these tools to risk measures, Wasserstein robustness, pricing and capital, portfolio drift, reserving cash-flow distributions, climate-prevention diagnostics, reinsurance, dependence uncertainty, capital allocation, distributional fairness diagnostics and longevity risk&#8230;<br></p></li><li><p><strong><a href="https://www.johndcook.com/blog/2026/07/03/does-additional-data-always-reduce-posterior-variance/">Does additional data always reduce posterior variance?</a></strong></p><p>Additional data does not always decrease the size of a confidence interval. This post will look at this from a Bayesian perspective. In general, new information reduces your uncertainty regarding whatever you&#8217;re estimating. The posterior distribution becomes more concentrated as more data are collected. That&#8217;s what happens &#8220;in general&#8221; but does it necessarily happen every time you get new data? Conceivably if you get surprising data, data that is very unlikely given your current prior, posterior uncertainty might increase&#8230;<br></p></li><li><p><strong><a href="https://blog.djnavarro.net/posts/2026-07-03_subscript/">Markdown styling in R plots and tables</a></strong><br>There is now a well-established &#8211; albeit informal &#8211; convention in R of using the &#8220;label&#8221; attribute as a way of storing natural language descriptions of variables, one that is supported by many packages for visualisation such as ggplot2 and some tabulation packages like table&#8230;I&#8217;m not entirely sure of the history behind this convention, but from what I can tell it seems to have emerged from the <em>haven</em> and <em>labelled</em> packages, which provide valuable tools built using this idea. For the purposes of this post I&#8217;ll keep it simple, and define some simple helper functions that make it easy to set, get and modify the labels associated with a data set&#8230;<code><br></code></p></li><li><p><strong><a href="https://www.ethanrosenthal.com/2016/07/20/lets-talk-or/">I&#8217;m all about ML, but let&#8217;s talk about OR</a><br></strong>There&#8217;s a better way! And I feel like nobody talks about it because the Data&#8217;s not Big, you&#8217;re not Learning Deep things, and there&#8217;s nary a chatbot in sight. It&#8217;s boring, old operations research, which was something that I guess your university offered, but nobody really knew what it meant. Full disclosure: I still don&#8217;t really know what it means. I do know that the job of the data scientist is to bring value to the company, and having some operations and optimization in your toolbelt is quite valuable!&#8230;<br></p></li><li><p><strong><a href="https://medium.com/@lz1955/samudra-2-a-fast-cheap-ai-ocean-model-now-at-the-scale-that-matters-b37883c62d51">Samudra 2: A Fast, Cheap AI Ocean Model, Now at the Scale That Matters</a><br></strong>M&#178;LInES&#8217; neural ocean emulator now runs multi-year simulations at eddy-permitting resolution on a single GPU, turning a supercomputer-scale job into one you can run hundreds of times over. That opens the door to faster, cheaper ocean information for shipping, fisheries, energy, insurance, and seasonal forecasting&#8230;<br></p></li><li><p><strong><a href="https://link.springer.com/article/10.1007/s40979-023-00146-z">Testing of detection tools for AI-generated text</a></strong></p><p>The study seeks to answer research questions about whether existing detection tools can reliably differentiate between human-written text and ChatGPT-generated text, and whether machine translation and content obfuscation techniques affect the detection of AI-generated text. The research covers 12 publicly available tools and two commercial systems (Turnitin and PlagiarismCheck) that are widely used in the academic setting. The researchers conclude that the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text. Furthermore, content obfuscation techniques significantly worsen the performance of tools..<br></p></li><li><p><strong><a href="https://patchwork.data-imaginist.com/">patchwork</a></strong><br><span>The goal of </span><code>patchwork</code><span> is to make it ridiculously simple to combine separate ggplots into the same graphic. As such it tries to solve the same problem as </span><code>gridExtra::grid.arrange()</code><span> and </span><code>cowplot::plot_grid</code><span> but using an API that incites exploration and iteration, and scales to arbitrarily complex layouts&#8230;</span><br></p></li><li><p><strong><a href="https://medium.com/@poojarysanket.03/knowledge-distillation-explained-part-1-how-large-models-teach-smaller-ones-0e77fc9f596a">Knowledge Distillation Explained, Part 1: How Large Models Teach Smaller Ones</a></strong></p><p>Much of a large model&#8217;s capability can be transferred to a significantly smaller model. This is the essence of <em>Knowledge Distillation</em> &#8212; a technique where a smaller student model learns from a larger teacher model, aiming to retain most of the teacher&#8217;s performance while being faster and cheaper to run&#8230;In Part 1 of this series, we trace the evolution of knowledge distillation from Geoffrey Hinton&#8217;s foundational 2015 paper through the Transformer era, covering DistilBERT, TinyBERT, and MiniLM&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1upl6er/managing_dealing_with_junior_data_scientists/">Managing/ Dealing with Junior Data Scientists? [Reddit]</a><br></strong>I&#8217;ve been in the &#8216;data science&#8217; space for a decade+ or so now. One thing I&#8217;ve noticed is that generally - give or take - outside of the elite jobs (&lt;2-3% aka not me and almost certainly not you) the caliber of coworkers has declined drastically. I&#8217;m not some fabled data scientist. I wasn&#8217;t some GitHub nerd who had everything embroil or terminal wizard nor could I write out the math to a GBM on a blackboard. I&#8217;d even forget basic obvious statistics. But I felt like I had common sense. Now I&#8217;m a manager/director. I work with data scientists. And I&#8217;m just generally freaked out by the absolute lack of basic common sense. This is across the last 7 that I have managed. Examples include&#8230;<br></p></li><li><p><strong><a href="https://www.johndcook.com/blog/2026/06/30/variables-and-parameters/">Distinguishing variables from parameters</a><br></strong>Imagine the following dialog.</p><p><strong>Professor</strong>: <em>f</em> is a function of a real variable <em>x</em> that takes a real parameter <em>k</em>.</p><p><strong>Student</strong>: What&#8217;s a parameter?</p><p><strong>Professor</strong>: It&#8217;s a constant that can vary.</p><p><strong>Student</strong>: Then if it can vary, isn&#8217;t it a variable?</p><p><strong>Professor</strong>: Sorta, but no not really.</p><p>This conversation plays out over and over, and unfortunately it often ends as it does above, with the student confused. Here&#8217;s how I believe the conversation should continue&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1uikmbi/what_is_the_most_underrated_skill_every_data/">What is the most underrated skill every data scientist should develop? [Reddit]</a></strong></p></li><li><p><strong><a href="https://moultano.wordpress.com/2026/06/19/where-to-find-the-colors-your-screen-cant-show-you/">Where to Find the Colors Your Screen Can&#8217;t Show You</a></strong></p></li><li><p><strong><a href="https://danunparsed.com/p/hackerrank-open-source-ats?hide_intro_popup=true">HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74/100. No &#8212; 88/100. Actually 83/100.</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** You can find last week's issue #658 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-658">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://solomonkurz.netlify.app/blog/2025-07-07-learn-stan-with-brms-part-i/">Learn Stan with brms, Part I</a></strong></p></li><li><p><strong><a href="https://raps-with-r.dev/">Building reproducible analytical pipelines with R</a></strong></p></li><li><p><strong><a href="https://data.post45.org/posts/small-press-distribution-bestsellers/">Small Press Distribution Bestseller Lists 2006-2023</a></strong></p></li></ul><p>.</p><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 658]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-658</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-658</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 02 Jul 2026 23:40:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WZbr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #659<br>June 26, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://moultano.wordpress.com/2026/06/19/where-to-find-the-colors-your-screen-cant-show-you/">Where to Find the Colors Your Screen Can&#8217;t Show You</a></strong><br>There are colors that I want to show you, but I can&#8217;t. They exist in the real world. You probably saw some of them today, but I can&#8217;t show them to you on a screen. A digital photograph can&#8217;t capture them, and your screen can&#8217;t display them. No game you&#8217;ve ever played has contained them. Unless you have specialized equipment, they are entirely absent from the digital world&#8230;<br></p></li></ul><ul><li><p><strong><a href="https://www.wespiser.com/posts/2026-06-19-best-dog-treat.html">Finding the Best Dog Treat with Statistics: A Greyhound, five treats, and a Bradley-Terry model</a></strong><br>Bebop, my 83lb, 33 inch tall, Greyhound, loves three things: running fast, following me around the house, and treats. Whether it&#8217;s a chew treat, pizza out of a child&#8217;s hand who strayed too far from a party, or a small tray of cat food, he has a nose for what he likes and the athleticism to give him a fair shot at getting it. I&#8217;ve watched him eat for years, so it was upsetting to realize I don&#8217;t know what his favorite snack is, and can&#8217;t easily ask him. Fortunately for Bebop&#8217;s palate, the Bradley-Terry model gives us a way to figure out a &#8220;strength&#8221; of treat from pairwise comparisons&#8230;</p><p></p></li><li><p><strong><a href="https://arxiv.org/abs/2602.03092">Generative Artificial Intelligence creates delicious, sustainable, and nutritious burgers</a></strong><br>Using burgers as a model system, the generative AI rediscovers the classic Big Mac without explicit supervision and generates novel burgers optimized for deliciousness, sustainability, or nutrition. Compared to the Big Mac, its delicious burgers score the same or better in overall liking, flavor, and texture in a blinded sensory evaluation conducted in a restaurant setting with 101 participants; its mushroom burger achieves an environmental impact score more than an order of magnitude lower; and its bean burger attains nearly twice the nutritional score. Together, these results establish generative AI as a quantitative framework for learning human taste and navigating complex trade-offs in principled food design&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:700536}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WZbr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WZbr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png 424w, https://substackcdn.com/image/fetch/$s_!WZbr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png 848w, https://substackcdn.com/image/fetch/$s_!WZbr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png 1272w, https://substackcdn.com/image/fetch/$s_!WZbr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WZbr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png" width="1140" height="762" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:762,&quot;width&quot;:1140,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:74138,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/204753762?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WZbr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png 424w, https://substackcdn.com/image/fetch/$s_!WZbr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png 848w, https://substackcdn.com/image/fetch/$s_!WZbr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png 1272w, https://substackcdn.com/image/fetch/$s_!WZbr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79318fb7-e52b-4c1e-b8ed-4e6e40818231_1140x762.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://danunparsed.com/p/hackerrank-open-source-ats?hide_intro_popup=true">HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74/100. No &#8212; 88/100. Actually 83/100</a></strong><br>How hiring is becoming a luck filter&#8230;Here a quick rundown on how the tool works: Your PDF gets parsed into text. An LLM is called six times to extract structured information &#8212; your basics, work history, education, skills, projects, awards. It pulls your GitHub profile, scans your top repos, appends them as extra context. Then everything gets fed into the LLM at once to be graded&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1uikmbi/what_is_the_most_underrated_skill_every_data/">What is the most underrated skill every data scientist should develop? [Reddit]</a></strong></p><p>Beyond Python, machine learning, and statistics, which skill has made the biggest difference in solving real-world data science problems and delivering business value?&#8230;</p><p></p></li><li><p><strong><a href="https://www.youtube.com/watch?v=cNpvQq5rJEs">Can AI Become a Real Data Scientist? | Ga&#235;l Varoquaux on scikit-learn, Probabl &amp; Scientific Judgment</a><br></strong>In this episode of Creative Difference, Maxime Gabella speaks with Ga&#235;l Varoquaux, co-founder of scikit-learn and Probabl, about the future of AI, statistical learning, and scientific judgment. They discuss the difference between generative AI and statistical machine learning, why current AI agents can write code but still struggle with rigorous data analysis, and why tools like scikit-learn and Skore may become essential for trustworthy AI systems. The conversation explores memorization, generalization, uncertainty, data leakage, evaluation, human-in-the-loop science, and the role of a &#8220;statistical harness&#8221; for AI agents. It also moves into deeper questions: why real-world science is harder than mathematics, why data collection and intervention still matter, how categories shape reality, and what creativity means for humans and machines&#8230;<br></p></li><li><p><strong><a href="https://capestart.com/resources/blog/vibe-coding-vs-vibe-engineering/">Vibe Coding vs. Vibe Engineering: How Systems Scale Without Collapsing</a><br></strong>What is the difference between vibe coding and vibe engineering? Learn how modern software teams evolve from rapid experimentation to enterprise-grade systems without drowning in technical debt or sliding into technical bankruptcy&#8230;<br></p></li><li><p><strong><a href="https://rpsychologist.com/descriptive-adjustment/">Why Adjusted Regression Coefficients Are Less Descriptive Than They Look</a></strong></p><p>It is common for researchers to investigate &#8220;factors associated with&#8221; an outcome by collecting many candidate variables, entering them together into one multivariable regression, and reporting the ones whose mutually-adjusted coefficients reach significance (Lewer et al., 2025). In this interactive article I explain why that practice misleads &#8212; not because of the usual problems of causal inference or multiple comparison, but even in descriptive studies, where researchers are genuinely interested in associations&#8230;<br></p></li><li><p><strong><a href="https://tensor4all.org/blog/introducing-tenferro-rs/">From Julia to Rust: a differentiable tensor stack for scientific computing in the agentic AI era</a></strong><br><span>The Rust ecosystem has changed a lot in the last few years. crates.io went from 602 crates in 2015 to roughly 210,000 in 2026 (</span><a href="https://github.com/shinaoka/rust_crate_count">data</a><span>). For dense linear algebra there is faer; for GPU kernels, </span><a href="https://github.com/tracel-ai/cubecl">CubeCL</a><span>; for generic numerics, </span><code>num-traits</code><span> and </span><code>num-complex</code><span>. There are also libraries at nearby layers: ndarray for arrays, nalgebra and faer for linear algebra, Burn and candle for deep learning, and numr for a NumPy-style array API. What we needed was the layer between them: a scientific-computing tensor stack with column-major storage, dynamic shapes, eager and traced autodiff, einsum, FFT, CPU/CUDA backends, and extensible operations. That is what tenferro-rs is for&#8230;</span>This post explains why we are building it, and why we chose Rust now that code is no longer written only by humans&#8230;<code><br></code></p></li><li><p><strong><a href="https://m.canouil.dev/gribouille/">Gribouille - Create elegant graphics with the Grammar of Graphics for Typst.</a><br></strong>Compose charts by layering data, aesthetic mappings, geoms, scales, and themes. The same grammar as ggplot2 and plotnine, drawn natively in Typst&#8230;<br></p></li><li><p><strong><a href="https://research.getrecast.com/geolift-sim-study/">Open-source geo-experiment tools are </a></strong><em><strong><a href="https://research.getrecast.com/geolift-sim-study/">not</a></strong></em><strong><a href="https://research.getrecast.com/geolift-sim-study/"> interchangeable</a><br></strong>We ran 32,000 simulated experiments across four common marketing scenarios to benchmark four leading open-source geo-experiment tools. Because we use synthetic data where the true campaign effect is known in advance, we can measure how well each tool recovers it. These four tools are often treated as interchangeable, and they are not. Where they diverge &#8212; sharply, at times &#8212; is in how they handle uncertainty: how often their confidence intervals contain the true incremental effect, how often they declare winning results that aren&#8217;t real (false positives), and how often they come back inconclusive when a real incremental effect exists (false negatives)&#8230;<br></p></li><li><p><strong><a href="https://clickhouse.com/blog/open-source-10">Ten years of ClickHouse in open source</a></strong></p><p><span>ClickHouse was released in open source on </span><a href="https://news.ycombinator.com/item?id=11908254">Jun 15 2016</a><span>, ten years ago. Since then, it became the most popular open source analytical database with more than 2000 contributors&#8230;</span><br></p></li><li><p><strong><a href="https://fergusfinn.com/blog/what-happens-when-you-run-a-gpu-kernel/">What happens when you run a CUDA kernel</a></strong><br>Here&#8217;s a simple CUDA program. It adds two vectors&#8230;Compiled for an RTX 4090, and launched, it does correctly work out that 1 + 1 = 2, a million times&#8230;Telling you that involved tens of millions of CPU instructions, a couple of device files, nine hundred ioctls, and one memory-mapped doorbell register. In this post, we&#8217;ll follow this one kernel from the code down to the warps, and back up to the answer&#8230;<br></p></li><li><p><strong><a href="https://blog.greg.technology/2026/06/12/map-clustering-is-not-my-favorite.html">Map Clustering is Not My Favorite</a></strong></p><p>Friends, this post has been in the making for maybe, 20 years. Google Maps (and me) are (at least) that old. And mapping, of course, existed before that; those were the days. But ~~~20 years ago, or like a year after that, something appeared on our collective mapping radars that I have not been able to shake off since - and which I haven&#8217;t kvetched about in writing, ever, for everyone to &#8220;enjoy&#8221;. Enjoy! It&#8217;s a Greg Rant! Amazing&#8230; What happened ~~~~20 years ago, after Google Maps was hoisted on us all, was that &#8220;people wanted to see more than 1 point on a map&#8221; and sometimes that meant 100 points, or a thousand! And things were not easy or good then. I mean, 1000 points, that&#8217;s a LOT. But then, who said 1000 points (was it you?) was too much to show? Well&#8230; browsers&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1u8r8yi/data_directors_whats_your_next_step/">Data Directors - what&#8217;s your next step? [Reddit]</a><br></strong>For anyone who has had a director of data or data director title in the past - where are you now? Similar role at a different company? Same role? Eventually C-suite? What&#8217;s the plan?&#8230;<br></p></li><li><p><strong><a href="https://blog.jxmo.io/p/zen-and-the-art-of-machine-learning?hide_intro_popup=true">Zen and the Art of AI Research (temperament &gt;&gt; talent)</a><br></strong>So you want to do AI research? It&#8217;s true that no one really teaches you how. Not directly, anyway. The way to get started is pretty simple: some combination of (i) reading and (ii) building stuff. You can&#8217;t do one without the other. You become a researcher through the combination. It turns out the process of becoming a great researcher is not unlike learning to meditate&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://yongzx.github.io/blog/2026/06/24/job-search/">Surprising lessons from my research scientist job search</a></strong></p></li><li><p><strong><a href="https://medium.com/operations-research-bit/the-pca-mistake-i-made-during-my-phd-and-how-to-avoid-it-in-r-4bbb148b2bb5">The PCA Mistake I Made During My PhD (and How to Avoid It in R)</a></strong></p></li><li><p><strong><a href="https://freerangestats.info/blog/2026/06/23/uk-prime-ministers">United Kingdom prime ministers</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** You can find last week's issue #657 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-657">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://www.sciencedirect.com/science/article/pii/S0045782526004445"><span>Generative AI for material design: A mechanics perspective from burgers to matter</span></a></strong></p></li><li><p><strong><a href="https://medium.com/@s_nikolaev/faster-knn-search-in-manticore-2-pass-hnsw-batched-distances-and-avx-512-b85604647aab">Faster KNN search in Manticore: 2-pass HNSW, batched distances, and AVX-512</a></strong></p></li><li><p><strong><a href="https://github.com/nmslib/hnswlib">Hnswlib - Header-only C++/python library for fast approximate nearest neighbors</a></strong></p></li><li><p><strong><a href="https://htmx.org/essays/working-with-ai/">Working With AI: A Concrete Example</a></strong></p></li><li><p><strong><a href="https://ayushtambde.com/blog/matrix-orthogonalization-improves-memory-in-recurrent-models/">Matrix Orthogonalization Improves Memory in Recurrent Models</a></strong></p></li><li><p><strong><a href="https://rozumem.xyz/posts/17">The Doorman Fallacy In Action</a></strong></p></li><li><p><strong><a href="https://www.greybeam.ai/blog/duckdb-internals-part-1">DuckDB Internals: Why is DuckDB Fast? (Part 1)</a></strong></p></li><li><p><strong><a href="https://kolistat.com/blog/the-stats-duck-v0-6-0/">the-stats-duck v0.6.0 &#8212; statistics that live in your SQL</a></strong></p></li><li><p><strong><a href="https://corinwagen.github.io/public/blog/20260701_information_content.html">Small Molecules Have More Information Per Atom Than Biologics</a></strong></p></li></ul><p>.</p><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 657]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-657</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-657</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 25 Jun 2026 22:51:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8-d4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #657<br>June 25, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://freerangestats.info/blog/2026/06/23/uk-prime-ministers">United Kingdom prime ministers</a><br></strong>UK has had a spurt of prime ministerial turnover in the past decade or so, but it&#8217;s by no means unprecedented. I download data from Wikipedia and try several ways to visualise that turnover&#8230;</p></li></ul><ul><li><p><strong><a href="https://perthirtysix.com/how-the-heck-do-synthesizers-work">How The Heck Do Synthesizers Work? (An Interactive Exploration)</a></strong><br>Synthesizers have remained a staple of modern music. From iconic video game soundtracks to Hans Zimmer&#8217;s scores, from Radiohead to the Stranger Things theme. We hear more synthesized music than we even realize. So how do they actually work?&#8230;</p><p></p></li><li><p><strong><a href="https://yongzx.github.io/blog/2026/06/24/job-search/">Surprising lessons from my research scientist job search</a></strong><br>There are two recent blog posts from Alisa and Silvia, both CS PhD students, on how they prepared and got into frontier labs such as OpenAI and Google Deepmind. I highly recommend them, and after seeing the reactions on Twitter, I want to share a different angle: what surprised me during my own research scientist job search&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:652703}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8-d4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8-d4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png 424w, https://substackcdn.com/image/fetch/$s_!8-d4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png 848w, https://substackcdn.com/image/fetch/$s_!8-d4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png 1272w, https://substackcdn.com/image/fetch/$s_!8-d4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8-d4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png" width="1140" height="718" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:718,&quot;width&quot;:1140,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:73383,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/203612785?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8-d4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png 424w, https://substackcdn.com/image/fetch/$s_!8-d4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png 848w, https://substackcdn.com/image/fetch/$s_!8-d4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png 1272w, https://substackcdn.com/image/fetch/$s_!8-d4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1203554e-3011-44d7-945c-f8e8083aa356_1140x718.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://medium.com/operations-research-bit/the-pca-mistake-i-made-during-my-phd-and-how-to-avoid-it-in-r-4bbb148b2bb5">The PCA Mistake I Made During My PhD (and How to Avoid It in R)</a></strong><br>A few years into my PhD, I ran a principal component analysis on a dataset I&#8217;d spent six months collecting, looked at the scree plot, picked &#8220;the number of components that looked right,&#8221; and moved on. A reviewer later asked me to justify that number. I couldn&#8217;t. Not with anything more rigorous than &#8220;the elbow looked like it was there.&#8221;&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/analytics/comments/1ud7d31/it_turns_out_analytics_was_a_great_career_to_go/">It turns out Analytics was a great career to go into even in a world with AI [Reddit]</a></strong></p><p>Maybe two or three years ago I lamented the fact I had never gone into software development in spite of the fact I probably had the coding mindset for it, regretting the tedious and stressful aspects of Analytics as well as lower overall pay. Now with AI leading to massive layoffs and / or reduce hiring in software development and other Engineering fields, I&#8217;m thinking Analytics was a good field to specialize in since it has that sweet spot of being just close enough to the business and just close enough to the tech side that it is hard to automate away via AI. Furthermore, I think demand for analysts in general to understand data and accommodate reporting changes will also increase if AI is accelerating software changes and changes to data models and systems&#8230;</p><p></p></li><li><p><strong><a href="https://github.com/purohit10saurabh/minFLUX">minFLUX - A hackable implementation of FLUX diffusion models</a><br></strong>A simplified educational PyTorch implementation of FLUX.1 and FLUX.2 diffusion transformers (DiT) by Black Forest Labs. Built for understanding rectified flow matching, joint attention, and the key design choices behind FLUX with verifiable line-by-line source mappings to the official codebases&#8230;<br></p></li><li><p><strong><a href="https://arxiv.org/abs/2509.06735">Data-driven discovery of dynamical models in biology</a><br></strong>In this review, we survey approaches for model discovery in biological dynamical systems, focusing on three methodological families: regression-based methods, network-based architectures, and decomposition techniques. We compare their ability to address three core goals: forecasting future states, identifying interactions, and characterizing system states. Representative methods are applied to a common benchmark, the Oregonator model, a minimal nonlinear oscillator that captures shared design principles of chemical and biological systems. By highlighting strengths, limitations, and interpretability, we aim to guide researchers in selecting tools for analyzing complex, nonlinear, and high-dimensional dynamics in the life sciences&#8230;<br></p></li><li><p><strong><a href="https://lilianweng.github.io/posts/2026-06-24-scaling-laws/">Scaling Laws, Carefully</a></strong></p><p>Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss <strong>L</strong> decreases predictably as we scale up model size <strong>N</strong>, dataset size <strong>D</strong>, and compute <strong>C</strong>, following a power-law curve, which appears as a straight line on a log-log plot. We can view scaling laws as a framework for describing the relationship between compute, loss, model size and data; at its core, it is about how to allocate precious compute optimally between <strong>N </strong>and <strong>D</strong>. This predictability makes scaling laws highly valuable in practice. A common workflow is to fit scaling laws on a handful of small runs and then extrapolate to estimate the token and compute requirements for larger models&#8230;.<br></p></li><li><p><strong><a href="https://www.youtube.com/watch?v=MMDNaeIFVy8&amp;t=3s">ML Foundations (prerequisites) for Post-Training | RLHF Book Course, Lecture 0</a></strong><br>In this video I try to cover a bunch of math, LLM training fundamentals, and probability concepts that come up again and again in post-training content (and this book). We cover things like the role of mid-training, definitions of KL, entropy &amp; cross-entropy, getting LM probabilities from a sequence, etc&#8230;<code><br></code></p></li><li><p><strong><a href="https://github.com/imann128/tsauditor">tsauditor - data quality auditing library for time-series tabular data in financial and sensor domains</a><br></strong>A data-quality auditing library for time-series tabular data, with a focus on financial and sensor domains. tsauditor scans a DataFrame and returns a structured report of structural problems, anomalies, and &#8212; its core contribution &#8212; data-leakage between features and the prediction target&#8230;<br></p></li><li><p><strong><a href="https://www.nature.com/articles/s41591-026-04431-5">General-purpose large language models outperform specialized clinical AI tools on medical benchmarks</a><br></strong>Specialized clinical artificial intelligence (AI) tools are entering medical practice at scale1,2. These proprietary large language model (LLM)-based tools promise superior clinical performance to general-purpose frontier LLMs as a result of domain-specific training or retrieval-augmented generation (RAG)3. Yet, their architectures, base models and training pipelines are not public&#8230;This study is an independent, quantitative comparison of clinical AI tools against frontier LLMs using real-world physician queries from the course of care. Clinical AI tools lagged behind frontier models on every evaluation: knowledge, expert alignment and real-world clinical use across multiple dimensions&#8230;<br></p></li><li><p><strong><a href="https://praxiscurrents.substack.com/p/moneyball-for-physical-ai">Moneyball for Physical AI - A Scaling Law Perspective for Marginal Utility per Dollar</a></strong></p><p>This essay builds a framework for the marginal utility of data, and uses it to discuss value accrual in Physical AI. We take the perspective of the scaling laws that guide how loss behaves with data, and the unit economics that govern what a dollar of data is worth. Together they give an approximate marginal utility per dollar, the on-base percentage of physical AI&#8230;<br></p></li><li><p><strong><a href="https://statswithcats.net/2026/06/20/five-eras-in-the-evolution-of-probability-and-statistics/">Five Eras in the Evolution of Probability and Statistics</a></strong><br>Statistics isn&#8217;t a new thing. It dates back at least fifty centuries beginning as counts in the form of tally marks for keeping track of crops, animals, people, and time. From there, it evolved with the demands of government and business, supported by academic inquiry and the growth of technology&#8230;<br></p></li><li><p><strong><a href="https://allmyfriendsarestories.neocities.org/AParchive/resources">Interested in learning more about digital archives?</a></strong></p><p>Interested in learning more about digital archives? Here are some resources to get you started! Most are targeted at a general audience, but some of them are more technical. I&#8217;ve included descriptions to help you navigate. Let me know if you have any topics you&#8217;d like me to cover here!..<br></p></li><li><p><strong><a href="https://www.reddit.com/r/analytics/comments/1ubkh7f/anyone_actually_believe_dashboards_are_going_away/">Anyone actually believe dashboards are going away? [Reddit]</a><br></strong>keep seeing this take that ai agents are going to replace dashboards entirely. like why even look at a chart when the ai can just tell you "stop spending on this channel" or whatever. and i get the appeal of that but i think it misses something pretty fundamental about how people actually make decisions&#8230;<br></p></li><li><p><strong><a href="https://www.seascapemodels.org/posts/2026-06-25-prediction-vs-fitting/">Data for fitting models versus data for predicting from models</a><br></strong>Answering a question that came up from a student recently. Say you have 20 surveys of reef fish biomass at different locations. Then you also have gridded data with environmental covariates. The gridded data is for all reefs everywhere. The goal is to predict fish biomass at all reefs everywhere. Here&#8217;s an older post that walks through the steps in R with older packages (you will want to update raster to terra, everything else should work). The statistically correct workflow would look like this...</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1u7upnm/2026_tech_stack_at_your_job/">2026 Tech Stack at your Job [Reddit]</a></strong></p></li><li><p><strong><a href="https://sqltoerdiagram.com/">Free SQL&#8594;ER diagram tool, runs in the browser, nothing uploaded</a></strong></p></li><li><p><strong><a href="https://vickiboykis.com/2026/06/15/running-local-models-is-good-now/">Running local models is good now</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** Please take a look at last week's issue #656 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-656">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://textaural.com/uva_HDC/playlist.html">Histories of Digital Culture</a></strong></p></li><li><p><strong><a href="https://unconv.ai/blog/introducing-un-0-generating-images-with-coupled-oscillators/">Introducing Un-0: Generating Images with Coupled Oscillators</a></strong></p></li></ul><p>.</p><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 656]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-656</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-656</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 18 Jun 2026 23:25:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!eFK_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #656<br>June 18, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://sqltoerdiagram.com/">Free SQL&#8594;ER diagram tool, runs in the browser, nothing uploaded</a><br></strong>Paste a SQL schema (CREATE TABLE statements) &#8594; get a clean, interactive ER diagram. Open source and 100% local &#8212; it runs entirely in your browser, so your schema never leaves your machine. No server, no signup, no upload&#8230;</p></li></ul><ul><li><p><strong><a href="https://osf.io/preprints/psyarxiv/qpj6n_v2">Preschoolers search semantic networks in a broader and more variable way than adults: Implications for hypothesis generation</a></strong><br>We find that adults show greater dependencies between sequential guesses than preschoolers, and generate a less diverse set of options. These findings may support the idea that development can be viewed as analogous to simulated annealing strategies in machine learning that start &#8220;hot&#8221; (in early childhood), generating wider and more variable searches, and eventually cool (in adulthood) to generate narrower searches&#8230;</p><p></p></li><li><p><strong><a href="https://rworks.dev/posts/too-many-R-packages/">New CRAN Packages: signal or noise?</a></strong><br>CRAN continues to be the most accessible repository for statistical knowledge on the planet, and the number of new packages being accepted by CRAN is growing faster than ever. But, is the R community really benefiting from this new growth?&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:612414}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eFK_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eFK_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png 424w, https://substackcdn.com/image/fetch/$s_!eFK_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png 848w, https://substackcdn.com/image/fetch/$s_!eFK_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png 1272w, https://substackcdn.com/image/fetch/$s_!eFK_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eFK_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png" width="633" height="394.7979094076655" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:716,&quot;width&quot;:1148,&quot;resizeWidth&quot;:633,&quot;bytes&quot;:81496,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/202651355?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eFK_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png 424w, https://substackcdn.com/image/fetch/$s_!eFK_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png 848w, https://substackcdn.com/image/fetch/$s_!eFK_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png 1272w, https://substackcdn.com/image/fetch/$s_!eFK_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bfa9a21-0ac9-4482-a3b4-0c9c507fe6f4_1148x716.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://shonczinner.github.io/posts/embedding-prediction/">The 90-year-old idea behind JEPA models: Canonical Correlation Analysis (CCA)</a></strong><br>Harold Hotelling&#8217;s 1936 Canonical Correlation Analysis (CCA) [modern terminology, &#8220;CCA is used to find a common signal among two large matrices&#8221;] forms the theoretical and intuitive foundation for modern embedding prediction techniques, including JEPA models&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1u7upnm/2026_tech_stack_at_your_job/">2026 Data Science Tech Stack at your Job [Reddit]</a></strong></p><p>What is your current tech stack at your job?</p><p></p></li><li><p><strong><a href="https://interlatent.com/blog/interlatent-robotics-hardware-guide">A Beginner&#8217;s Guide to Robotics Hardware</a><br></strong>Building an open-source robot begins in a way that is familiar to assembling IKEA furniture&#8230;The similarity ends once the robot is powered on. A bookshelf is designed to stay exactly as assembled, whereas a robot has to move while remaining correct about its own position and the state of its surroundings&#8230;When designing and building robots, this need for correctness both numerically and temporally is a key consideration&#8230;The rest of this post looks at how that difference manifests, using a common framing in robotics that divides the hardware into three parts: the movement, the body, and the sensor&#8230;<br></p></li><li><p><strong><a href="https://www.youtube.com/playlist?list=PLGVZCDnMOq0rFQykYJg7t441AEpN4SszE">PyData London 2026 Talks</a><br></strong>All the talks from the PyData London 2026 are now available&#8230;<br></p></li><li><p><strong><a href="https://www.ssp.sh/brain/data-engineering-acquisitions/">Data Engineering Acquisitions (2022-2026)</a></strong></p><p>Consolidation in the Data Engineering market is happening quickly. Tools from the Modern Data Stack get unified into bigger Data Platforms. This note highlights the latest acquisitions across data engineering. It serves as an overview of the latest consolidations. Find attached the acquisition overview from 2022 to today&#8230;<br></p></li><li><p><strong><a href="https://apenwarr.ca/log/20260531">The software industry: annealing, but wrong</a></strong><br>In recent months I&#8217;ve heard of several teams with an interesting policy: each pull request should be no more than a few files, and no more than a certain number of lines (say 500). And do just one thing and do it well. And be easy for a human to review. And be fully tested by the test suite&#8230;And often, the results are good. Sure, splitting a single 6000-line feature or fix into twelve 500-line PRs is more work, but each of those PRs is surely easier to review. And you can git bisect them when there&#8217;s a bug! And maybe revert the individual change that broke something. ...and also cause 12x as many context switches for your reviewers as they review each one sequentially. But that&#8217;s just the cost of software quality! Right?&#8230;<code><br></code></p></li><li><p><strong><a href="https://www.sharonlohr.com/blog/2026/6/10/ai-and-sampling-problems">AI and Survey Sampling Problems</a><br></strong>My previous post discussed the performance of the artificial intelligence (AI) interface Gemini on undergraduate statistics problems. Now let&#8217;s look at how Gemini answers some of the problems in my sampling textbook (Lohr, 2022), and talk about how Gemini could help students learn sampling&#8230;<br></p></li><li><p><strong><a href="https://jakubsobolewski.com/blog/test-doubles-taxonomy/">Test Doubles Taxonomy for R: Dummy, Stub, Spy, Mock, Fake</a><br></strong>You might call them all &#8220;mock&#8221;&#8230;Mock the database. Mock the API. Mock the function. The word becomes a catch-all for any test double, any object you substitute for a real dependency in a test. Lumping them together makes it harder to choose the right tool, and the wrong choice leads to brittle, misleading tests. There are five distinct types, each with a specific job. Knowing which is which is how you stop writing tests that do the wrong thing&#8230;<br></p></li><li><p><strong><a href="https://iquilezles.org/articles/fbm/">Fractional Brownian Motion</a></strong></p><p>A Brownian Motion (BM), without the &#8220;fractional&#8221; part, is a motion where the position of a given object over time changes in random increments&#8230;A Fractional Brownian Motion is a similar process in which the increments are not completely independent from each other, but there&#8217;s some sort of memory to the process&#8230;I believe fBM is not necessarily a well understood mechanism. So this article describes how it functions and their different spectral and visual characteristics for various values of their main parameter H, backed with some experiments and measurements&#8230;<br></p></li><li><p><strong><a href="https://vickiboykis.com/2026/06/15/running-local-models-is-good-now/">Running local models is good now</a></strong><br>I&#8217;ve been working with local models since they came out, and finally, they&#8217;re surprisingly good now&#8230;<br></p></li><li><p><strong><a href="https://jakubsobolewski.com/blog/snapshot-testing-beyond-screenshots/">Snapshot Testing in R: Beyond Screenshots</a></strong></p><p>In this post I want to walk through using snapshot testing for what it is good for, and the practices that make it efficient&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/AskProgramming/comments/1u5lbig/what_algorithm_is_surprisingly_new/">What algorithm is surprisingly new? [Reddit]</a><br></strong>Other than any AI stuff, I&#8217;m talking about the types of algorithms you learn about in any standard Data Structures and Algorithms University course&#8230;I&#8217;m surprised that alot of these algorithms were actually invented HUNDREDS of years ago&#8230;<br></p></li><li><p><strong><a href="https://arxiv.org/abs/2606.02184">The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing</a><br></strong>These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived. We show that llms do not merely default to high-probability individual names when generating fictional experts: they produce correlated character ensembles, pairs and trios whose co-occurrence rates far exceed chance and are consistent across independent generations. These priors are model-family-specific (Claude: Elena Vasquez + Marcus Chen + Amara Okafor; Gemini: Aris Thorne + Lena Petrova; GPT: Elara Voss with no fixed partner)&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://otexts.com/fpppy/">Forecasting: Principles and Practice, the Pythonic Way</a></strong></p></li><li><p><strong><a href="https://maxhalford.github.io/blog/solution-engineering-advice/">My unvarnished guide to solution engineering</a></strong></p></li><li><p><strong><a href="https://interlatent.com/blog/interlatent-modern-ai-robotics-first-principles">An Overview of Modern AI Robotics from First Principles</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** Please take a look at last week's issue #655 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-655">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://r-posts.com/2026-rousseeuw-prize-for-statistics-awarded-to-r-core-team-for-transforming-statistics-computing-worldwide/">2026 Rousseeuw Prize for Statistics Awarded to R Core Team for Transforming Statistics Computing Worldwide</a></strong></p></li><li><p><strong><a href="https://www.kenkoonwong.com/blog/aa1/">Learning Amino Acids Part 1: Non-Polar Amino Acids, Rodrigues Rotation, and Lennard-Jones Potential</a></strong></p></li><li><p><strong><a href="https://statmodeling.stat.columbia.edu/2026/06/18/gambling-provides-a-gentle-rocking-of-the-emotions-to-put-you-in-a-pleasant-baby-like-state/">Gambling provides a gentle rocking of the emotions to put you in a pleasant baby-like state</a></strong></p></li><li><p><strong><a href="https://statswithcats.net/2026/06/12/the-caitlin-clark-effect/">The Caitlin Clark Effect</a></strong></p></li><li><p><strong><a href="https://blog.janestreet.com/formal-methods-at-jane-street-index/">Formal methods and the future of programming</a></strong></p></li><li><p><strong><a href="https://john.soban.ski/polars2.html">Crunch Big Data on Your Laptop With Polars Streaming</a></strong></p></li></ul><p>.</p><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 655]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-655</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-655</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 11 Jun 2026 22:25:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FVg8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #655<br>June 11, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://interlatent.com/blog/interlatent-modern-ai-robotics-first-principles">An Overview of Modern AI Robotics from First Principles</a><br></strong>There is a deceptively simple way to describe what physical AI is all about, a way in which anyone with a STEM background will intuitively understand. Like all other AI models, a model which controls a robot is also a function. It takes in observations (camera pixels, joint angles, the felt resistance of a gripper, etc) and it outputs actions, the next set of positions and torques for its motors&#8230;If you&#8217;ve ever trained a model that maps inputs to outputs, you can already grasp the shape of the problem. The interesting part is what happens when you take this familiar shape and drop it into a moving, active world&#8230;This sounds like ordinary machine learning, and for a while you can pretend it is. But robotics introduces a third axis that classic ML never had to respect: inference time&#8230;</p></li></ul><ul><li><p><strong><a href="https://maxhalford.github.io/blog/solution-engineering-advice/">My unvarnished guide to solution engineering</a></strong><br>Nowadays I feel more or less comfortable interacting with customers. But I was awful at first. I know because one of the cofounders gave me harsh feedback after a call with our first serious customer. I still remember slamming the lid of my computer when we debriefed. What I perceived as harsh feedback at the time turned out to help me grow quickly&#8230;I used to be a regular data scientist assigned to internal projects. Talking to prospects and customers got me out of my comfort zone. You owe them a service, and they expect you to deliver something. If something goes wrong they&#8217;ll go above your head to your founders, at which point you start feeling the heat. It can be quite harsh. But it can also be rewarding when things go well&#8230;</p><p></p></li><li><p><strong><a href="https://myzopotamia.dev/navier-stokes-fluid-simulation-explained-with-godot">Navier-Stokes fluid simulation explained with Godot game engine</a></strong><br>Let me start with the mathematical description of what we will do in this blog post. This description might sound daunting, but don&#8217;t worry - we&#8217;ll explain everything as we go. Here goes: we will simulate fluid flow by moving a scalar density field through a vector velocity field. We&#8217;ll simulate velocity diffusion and advection as well as density diffusion and advection. Then we will add velocity projection with the goal of making the fluid obey the law of mass conservation - which will happen by balancing divergence with a pressure field. We will use bilinear interpolation and Gauss-Seidel relaxation for approximating values where needed&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:572144}" data-component-name="PollToDOM"></div><p></p><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FVg8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FVg8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png 424w, https://substackcdn.com/image/fetch/$s_!FVg8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png 848w, https://substackcdn.com/image/fetch/$s_!FVg8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png 1272w, https://substackcdn.com/image/fetch/$s_!FVg8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FVg8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png" width="596" height="367.8111888111888" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:706,&quot;width&quot;:1144,&quot;resizeWidth&quot;:596,&quot;bytes&quot;:69118,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/201616122?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FVg8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png 424w, https://substackcdn.com/image/fetch/$s_!FVg8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png 848w, https://substackcdn.com/image/fetch/$s_!FVg8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png 1272w, https://substackcdn.com/image/fetch/$s_!FVg8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a27a4d4-65ab-4318-8de5-ec2de8e962f7_1144x706.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://liangchang.substack.com/p/the-anti-scaling-law-in-biology-and">The Anti-Scaling Law in Biology, and Why AI Could Make Crowding Worse Before Making Drug Development Better</a></strong><br>One of the main reasons for the tech community&#8217;s optimism is the scaling-law. Once you demonstrated 0-1, you can do 1-100 much quicker. The internet, social media, and so on&#8230;In biology and drug development, I think there is a mirror image, the anti-scaling law. Because of that, here&#8217;s my contrarian view: AI could make crowding in drug development worse, before making it better. And that&#8217;s my perspective as a genuine believer in the transformative power of AI, and an AI practitioner who used $14,000 worths of AI tokens in the past 2 months&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/statistics/comments/1u1xnl1/what_is_there_besides_frequentist_and_bayesian/">What is there besides Frequentist and Bayesian stats? [Reddit]</a></strong></p><p>I am wondering whether there are lesser known statistical paradigms. like most people, I was first acquainted with the Frequentist framework, and later got introduced to Bayesian stats. I really like the way this made me reconsider some of what I thought were basic assumptions, so now I&#8217;m wondering what the next thing could be? Are there any other branches/frameworks which are not as well known?&#8230;</p><p></p></li><li><p><strong><a href="https://otexts.com/fpppy/">Forecasting: Principles and Practice, the Pythonic Way</a><br></strong>This textbook is based on Forecasting: Principles and Practice (3rd ed) and is intended to provide a comprehensive introduction to forecasting methods and to present just enough information about each method for readers to be able to use them sensibly. We don&#8217;t attempt to give a thorough discussion of the theoretical details behind each method, although the references at the end of each chapter will hopefully fill in many of those details&#8230;<br></p></li><li><p><strong><a href="https://medium.com/@VictorBanev/the-simplest-learning-machine-pt-2-e735367f546">The Simplest Learning Machine, Pt.2</a><br></strong>In the previous article I outlined the concept of the Simplest Learning Machine. It&#8217;s an imaginary algorithm that uses one byte of persistent memory and learns to predict something about a stream of binary events&#8230;Can we actually write something like that? How would it work?&#8230;One semi-obvious thing we can learn is the rate of positive events in the stream. This would give us some predictive power, as long as that rate is different from 50%. A bit of a stretch to call this &#8220;machine learning&#8221;, sure, but I&#8217;ll get to the questions of usefulness later&#8230;<br></p></li><li><p><strong><a href="https://research.dimensioncap.com/p/on-training-data-for-bio-ai-models">On Training Data for Bio AI Models</a></strong></p><p>As we advance biological foundation models, which lessons from LLM data curation transfer, and which need rethinking?&#8230;<br></p></li><li><p><strong><a href="https://ropensci.org/blog/2026/06/01/goodpractice/">Our goodpractice Package Has New Superpowers</a></strong><br>The goodpractice package has been recommended by rOpenSci since it was first started just over 10 years ago by G&#225;bor Cs&#225;rdi. We used to ask our editors to manually run goodpractice on all packages submitted to software peer-review, and then to ask authors to fix any notable issues flagged by the package&#8230;We&#8217;re really pleased to share that we&#8217;ve recently rolled out a host of updates and extensions to the package. These make it both easier to use, and more powerful&#8230;I think that is the aspect that I found most surprising: That the use of Claude made our collaboration feel less technical, and therefore somehow even more human. And that gave us the ability to work though 70 pull requests representing over 100 new checks, all ready for everybody to use&#8230;<code><br></code></p></li><li><p><strong><a href="https://rwarehouse.netlify.app/">The Warehouse</a><br>The Problem</strong></p><p>With over 23,000 packages on CRAN alone, finding the right package for your task is overwhelming:</p><ul><li><p>Searching by keywords often misses relevant packages</p></li><li><p>No easy way to compare similar packages</p></li><li><p>Quality indicators are scattered or missing</p></li><li><p>GitHub-only packages are hard to discover</p></li></ul><p></p><p><strong>The Warehouse Solution provides:</strong></p><ul><li><p><strong>Function-first search</strong>: &#8220;estimate serial interval&#8221; &#8594; find all relevant packages</p></li><li><p><strong>Quality scores</strong>: Automated assessment of tests, documentation, and maintenance</p></li><li><p><strong>All sources</strong>: CRAN, GitHub, Bioconductor in one place</p></li><li><p><strong>Community reviews</strong>: Real user experiences and recommendations</p></li><li><p><strong>Smart categorization</strong>: Browse by what packages actually do&#8230;<br></p></li></ul></li><li><p><strong><a href="https://cvg.ethz.ch/lectures/Robot-Learning/">Robot Learning: From Fundamentals to Foundation Models</a><br></strong>This course provides a comprehensive introduction to modern robot learning, combining classical techniques with the latest advances in large-scale models: Students will start by learning the fundamentals of imitation learning, reinforcement learning, and policy optimization, and gradually progress to advanced topics including Vision-Language-Action (VLA) models and foundation models for robotics The objectives of this course are:</p><ul><li><p>Understand the core principles of imitation learning, reinforcement learning, and policy learning.</p></li><li><p>Implement basic robot learning systems in simulation and on real robots.</p></li><li><p>Explore state-of-the-art Vision-Language Action and foundation models for robotics.</p></li><li><p>Design and evaluate scalable robot learning pipelines integrating perception, control, and multi-modal reasoning&#8230;<br></p></li></ul></li><li><p><strong><a href="https://silviasapora.github.io/blog/ml-interviews.html">ML Job Interviews: The Ultimate Guide</a></strong></p><p>How I found a Research Scientist role after a PhD in Machine Learning&#8230;My process was, overall, successful: I received offers from every company I completed interviews with including: DeepMind (which I accepted), Isomorphic Labs, Cohere, Meta, and a startup in stealth. A few caveats to the first claim: Anthropic, Mistral, and TeslaAI got back to me too late and I didn&#8217;t complete those processes. ReflectionAI, the one genuine rejection: they didn&#8217;t like me for the RS role but switched me to their Engineering track instead&#8230;<br></p></li><li><p><strong><a href="https://github.com/lucasduthu/stata-mpl">stata-mpl - Give your matplotlib and seaborn charts the Stata 19 look</a></strong><br>Give your matplotlib and seaborn charts the look of Stata 19 (the stcolor scheme, Stata&#8217;s colorblind-friendly default). Calibrated against the official SVG files exported by Stata 18/19&#8230;<br></p></li><li><p><strong><a href="https://theodore.net/projects/AvianVisitors/">Avian Visitors</a></strong></p><p>I mounted a tiny microphone on my apartment balcony to listen for any birds passing by and built a site to collage them as they&#8217;re heard&#8230;so I&#8217;ve thrown together this short writeup for any of you who want to monitor any avian visitors that may be passing by your own place. It&#8217;s short and sweet for now in an attempt to get something out quickly, but this work is part of a longer chain of bird-tangent projects i&#8217;ll write something up about soon!&#8230;<br></p></li><li><p><strong><a href="https://news.ycombinator.com/item?id=48449187">Ask HN: What are tools you have made for yourself since the advent of AI?</a><br></strong>Ask HN: What are tools you have made for yourself since the advent of AI?&#8230;<br></p></li><li><p><strong><a href="https://www.bayesianspectacles.org/why-academics-should-use-ai-for-writing-a-case-study/">Why Academics Should Use AI for Writing: A Case Study</a><br></strong>I violently dislike the idea of AI taking over my writing. My writing is my own, and having it done by AI makes the final product lose its soul. Also, whenever I have used AI to write several paragraphs independently (which I admit to doing for bureaucratic tasks) I ended up rewriting most of it anyway. However, over the past year or so I have become increasingly impressed with what AI can do, and rather than talk about this in abstract terms I would like to present you with a concrete demonstration that shifted my opinion a great deal&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://practicaldatacommunity.substack.com/p/how-to-build-a-simple-bulletproof">How to Build a Simple, Bulletproof Data Pipeline</a></strong></p></li><li><p><strong><a href="https://github.com/mljar/supertree">supertree - Interactive Decision Tree Visualization</a></strong></p></li><li><p><strong><a href="https://www.kaelio.com/blog/building-a-context-layer-for-the-agentic-era">Beyond the Semantic Layer: Building a Context Layer for the Agentic Era</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** Please take a look at last week's issue #654 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-654">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://statswithcats.net/2026/06/06/discover-stats-with-kittens-and-stats-with-cats/">When is detecting AI-generated text worthwhile?</a></strong></p></li><li><p><strong><a href="https://curlewis.co.nz/posts/lines-of-code-got-a-better-publicist/">Lines of Code Got a Better Publicist</a></strong></p></li><li><p><strong><a href="https://www.jstatsoft.org/article/view/v116i03">BayesMultiMode: Bayesian Mode Inference in R</a></strong></p></li><li><p><strong><a href="https://academic.oup.com/qje/article/140/2/943/7925870">Cognitive Endurance as Human Capital</a></strong></p></li><li><p><strong><a href="https://pub.sakana.ai/diffusionblocks/">DiffusionBlocks: Training Neural Networks One Block at a Time</a></strong></p></li><li><p><strong><a href="https://old.reddit.com/r/piano/comments/1u0lyhe/data_from_66000_practice_sessions_how_much_does/">Data from 66,000+ practice sessions. How much does the typical musician actually practice? [Reddit]</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1txc9vv/what_is_the_most_common_reason_data_science/?utm_source=share&amp;utm_medium=mweb3x&amp;utm_name=mweb3xcss&amp;utm_term=1&amp;utm_content=share_button">What is the most common reason data science projects fail to deliver business value? [Reddit]</a></strong></p></li></ul><p>.</p><div><hr></div><h2><strong>Whenever you're ready, 3 ways we can help:</strong><br></h2><ol><li><p><strong>Go deeper each week (paid subscription)</strong><br>Get 3 additional posts per week designed to help you:</p><ul><li><p>Statistics &#8594; understand the math behind ML</p></li><li><p>AI Agents &#8594; build with modern AI tools</p></li><li><p>Career &#8594; become more valuable at your job</p></li></ul><p><strong>&#128073; <a href="https://datascienceweekly.substack.com/subscribe">Upgrade for $10/month &#8212; cancel anytime</a><br></strong></p></li><li><p><strong>Looking to get a job?</strong><br>A practical guide to landing your first (or next) data science role, based on thousands of reader questions.<br><strong>&#128073; <a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">Check out our </a></strong><em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">&#8220;Get A Data Science Job&#8221;</a></strong></em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide"> Course</a></strong><br></p></li><li><p><strong>Promote your organization/project/event to ~68,500 subscribers<br></strong>Sponsor this newsletter and reach a highly engaged data science audience (30&#8211;35% open rate).<br><strong>&#128073; Reply to this email to learn more</strong></p></li></ol><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 654]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-654</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-654</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 04 Jun 2026 13:02:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rM6T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #654<br>June 04, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://practicaldatacommunity.substack.com/p/how-to-build-a-simple-bulletproof">How to Build a Simple, Bulletproof Data Pipeline</a><br></strong>In most organizations, a lot of value can be delivered with a setup that avoids high costs and unnecessary complexity. The most common scenario is not real-time streaming or exotic architectures. It is daily extraction from one or more transactional systems backed by a relational database&#8230;In this article, I want to walk through a simple but realistic example and show how a few design decisions, even when they look basic, can make a meaningful difference in robustness, operability, and long-term maintainability&#8230;</p></li></ul><ul><li><p><strong><a href="https://ankitg.me/blog/2026/05/04/fuzzy_api.html">AI for Bio has a Fuzzy API problem</a></strong><br>&#8220;AI for bio&#8221; is getting hot again. Given the excitement in the current moment, I thought I&#8217;d share a bit about what actually makes biology uniquely hard as an application domain for machine learning. The reason is not simply that biology is complicated, though it obviously is. ML is good at many things that are complicated. The deeper reason is that drug discovery does not have the kind of clean feedback loops and clean interfaces that made modern ML so powerful elsewhere&#8230;</p><p></p></li><li><p><strong><a href="https://momentsingraphics.de/Siggraph2026.html">Gaussian Point Splatting</a></strong><br>We propose Gaussian point splatting, a stochastic method to render Gaussian splats that scales extremely well to scenes with many Gaussians. Our core idea is to sample pixel-sized, opaque points from the Gaussians and to splat them to a framebuffer using 64-bit atomics. Through parallel programming primitives, we achieve an even distribution of the workload across millions of threads&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:529508}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rM6T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rM6T!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png 424w, https://substackcdn.com/image/fetch/$s_!rM6T!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png 848w, https://substackcdn.com/image/fetch/$s_!rM6T!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png 1272w, https://substackcdn.com/image/fetch/$s_!rM6T!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rM6T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png" width="1148" height="700" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:700,&quot;width&quot;:1148,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:65745,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/200601401?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rM6T!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png 424w, https://substackcdn.com/image/fetch/$s_!rM6T!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png 848w, https://substackcdn.com/image/fetch/$s_!rM6T!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png 1272w, https://substackcdn.com/image/fetch/$s_!rM6T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87268243-9c0b-410e-802b-a4d529bbb84a_1148x700.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://apps.london.gov.uk/state-of-london/">State of London Report, 2026</a></strong><br>The State of London report brings together an array of data about how London is performing across its economy, society and environment. Published annually by the GLA&#8217;s City Intelligence unit, it provides an up&#8209;to&#8209;date, data&#8209;led picture of life in the capital, drawing on a wide range of official and administrative sources. The report highlights trends and patterns and offers a shared evidence base to support debate and decision making&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/rstats/comments/1tqi2e3/what_is_considered_basic_r/">What is considered basic R? [Reddit]</a></strong></p><p>I have a job interview coming up and they want someone who knows basic R, I think I do have it, but what is your opinion on what it entails?&#8230;</p><p></p></li><li><p><strong><a href="https://humynlabs.ai/bridge">BRIDGE (Benchmark of Regional &amp; International Data for Global Evaluation) is the first independent Global South ASR benchmark</a><br></strong>BRIDGE (Benchmark of Regional &amp; International Data for Global Evaluation) is the first independent Global South ASR benchmark evaluating 15 global models across 22 languages on a first-of-its-kind 7 metric stack&#8230;<br></p></li><li><p><strong><a href="https://github.com/mljar/supertree">supertree - Interactive Decision Tree Visualization</a><br></strong><code>supertree</code> is a Python package designed to visualize decision trees in an interactive and user-friendly way within Jupyter Notebooks, Jupyter Lab, Google Colab, and any other notebooks that support HTML rendering. With this tool, you can not only display decision trees, but also interact with them directly within your notebook environment. Key features include:</p><ul><li><p>ability to zoom and pan through large trees,</p></li><li><p>collapse and expand selected nodes,</p></li><li><p>explore the structure of the tree in an intuitive and visually appealing manner&#8230;<br></p></li></ul></li><li><p><strong><a href="https://arxiv.org/abs/2606.02113">A Primer in Post-Training Reasoning Data: What We Know About How It Works</a></strong></p><p>Work on post-training reasoning data has grown rapidly, yet this literature remains scattered across dataset papers, reinforcement-learning recipes, reward-model studies, benchmarks, and frontier system reports. This paper is the first primer to synthesize over 150 key public studies and system reports on post-training reasoning data. We organize the field around four questions: what data objects exist, what makes them useful, how they are constructed, and how they scale. Together, this organization provides an attribution framework for future reasoning-data releases and post-training recipes&#8230;<br></p></li><li><p><strong><a href="https://christopherkrapu.com/blog/2026/dont-know-where-your-data-is-from/">Don&#8217;t know where your data is from? Bayesian modeling for unknown coordinates</a></strong><br>An especially strong motivating case for the usage of spatial probability models comes from the mining industry. During exploration for mineral resources, prospectors will take geologic samples by drilling holes and examining the resulting material for presence or concentration of valuable ores. These data typically show strong spatial correlation, but constructing a fully-detailed geophysical model is at times infeasible as we are able to observe very little of the underground conditions, though the advent of remote sensing techniques like ground-penetrating radar and gravimetry has dramatically improved our ability to characterize Earth&#8217;s subsurface. To address this challenge, we would like to construct a probability model which uses nearby data to predict a variable of interest at a new location&#8230;<code><br></code></p></li><li><p><strong><a href="https://blog.andymasley.com/p/why-i-think-panic-about-local-impacts?hide_intro_popup=true">Why I think panic about local impacts of data centers is just a panic</a><br></strong>In the last year of following increased local resistance to data centers being built, I&#8217;ve listened to lots of recorded testimony at town halls, read through countless comments and articles where people argue for why data centers are so uniquely evil and shut down anyone defending them as shills for AI companies&#8230;whenever I look into where people are actually getting their ideas about the hundreds of other data centers being built, the source always leads back to some confused misreading of local reporting, a wild calculation error, a bad game of telephone, or a wildly misleading article&#8230;<br></p></li><li><p><strong><a href="https://www.kaelio.com/blog/building-a-context-layer-for-the-agentic-era">Beyond the Semantic Layer: Building a Context Layer for the Agentic Era</a><br></strong>A context layer puts your warehouse schema, joins, metric definitions, and business knowledge in one reviewable place so data agents query governed context instead of guessing field names. A look at how it works, and at ktx, the open-source context layer&#8230;<br></p></li><li><p><strong><a href="https://lospino.so/statistics/jensen-shannon-divergence/">Jensen-Shannon Divergence</a></strong></p><p>How different are two discrete or binned probability distributions on the same support?&#8230;<br></p></li><li><p><strong><a href="https://www.mikeash.com/pyblog/fluid-simulation-for-dummies.html">Fluid Simulation for Dummies</a></strong><br>I wrote my Master&#8217;s thesis on high-performance real-time 3D fluid simulation and volumetric rendering. The basics of the fluid simulation that I used are straightforward, but I had a very difficult time understanding it. The available reference materials were all very good, but they were a bit too physics-y and math-y for me. Unable to find something geared towards somebody of my mindset, I&#8217;d like to write the page I wish I&#8217;d had a year ago. With that goal in mind, I&#8217;m going to show you how to do simple 3D fluid simulation, step-by-step, with as much emphasis on the actual programming as possible&#8230;.<br></p></li><li><p><strong><a href="https://www.sei.cmu.edu/blog/a-hitchhikers-guide-to-ml-training-infrastructure/">A Hitchhiker&#8217;s Guide to ML Training Infrastructure</a></strong></p><p>Hardware has made a huge impact on the field of machine learning (ML). Many of the ideas we use today were published decades ago, but the cost to run them and the data necessary were too expensive, making them impractical. Recent advances, including the introduction of graphics processing units (GPUs), are making some of those ideas a reality. In this post we&#8217;ll look at some of the hardware factors that impact training artificial intelligence (AI) systems, and we&#8217;ll walk through an example ML workflow&#8230;.<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1tqbfmq/weaponized_phrases_in_data_science_teams/">Weaponized phrases in Data science Teams [Reddit]</a><br></strong>&#8220;No free cycles&#8221; / &#8220;Empty plates&#8221;&#8230;&#8220;We need to focus on the low-hanging fruit&#8221;&#8230;"Be a go-getter, don't get stuck"&#8230;"Let's optimize our sprint velocity"&#8230;&#8220;You&#8217;re making this more complicated than it is&#8221;&#8230;"We need to relentlessly prioritize"&#8230;"I need you to own this initiative"&#8230;"Let's take this offline" / "Parking lot this"&#8230;"We need to leverage AI to unlock enterprise value"&#8230;"We're like a family here"&#8230;<br></p></li><li><p><strong><a href="https://taylorgeospatial.org/agricultural-field-boundaries-mapped-globally-for-the-first-time/">Agricultural Field Boundaries, Mapped Globally for the First Time</a><br></strong>For the first time, every agricultural field on Earth has a boundary on the map. Taylor Geospatial funded and co-developed this work with Microsoft AI for Good Lab because we believe GeoAI should work everywhere, not just in the data-rich regions where labeled training data is abundant. Today, it&#8217;s publicly available for everyone to benefit from&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1tknjuv/what_ds_job_market_trends_are_you_seeing/">What DS job market trends are you seeing? [Reddit]</a></strong></p></li><li><p><strong><a href="https://aiweekender.substack.com/p/6-llm-prompting-techniques-for-data">6 LLM Prompting Techniques for Data Scientists and Engineers in 2026</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/selfimprovement/comments/1tjef6k/i_am_faking_my_way_through_a_data_analyst_role/">I am faking my way through a Data Analyst role with AI, how do I actually learn before I get caught? [Reddit]</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** Please take a look at last week's issue #653 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-653">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://www.moderndescartes.com/essays/ai_and_expertise/">Expertise in the Age of AI</a></strong></p></li><li><p><strong><a href="https://vickiboykis.com/2026/05/28/we-should-be-more-tired-than-the-model/">We should be more tired than the model</a></strong></p></li><li><p><strong><a href="https://www.brethorsting.com/blog/2026/05/domain-expertise-has-always-been-the-real-moat/">Domain Expertise Has Always Been the Real Moat</a></strong></p></li><li><p><strong><a href="https://obeli.sk/blog/sqlite-is-all-you-need-for-durable-workflows/">SQLite is All You Need for Durable Workflows</a></strong></p></li><li><p><strong><a href="https://www.fharrell.com/talk/bguide/">Implications of the Draft FDA Bayesian Guidance</a></strong></p></li></ul><p>.</p><div><hr></div><h2><strong>Whenever you're ready, 3 ways we can help:</strong><br></h2><ol><li><p><strong>Go deeper each week (paid subscription)</strong><br>Get 3 additional posts per week designed to help you:</p><ul><li><p>Statistics &#8594; understand the math behind ML</p></li><li><p>AI Agents &#8594; build with modern AI tools</p></li><li><p>Career &#8594; become more valuable at your job</p></li></ul><p><strong>&#128073; <a href="https://datascienceweekly.substack.com/subscribe">Upgrade for $10/month &#8212; cancel anytime</a><br></strong></p></li><li><p><strong>Looking to get a job?</strong><br>A practical guide to landing your first (or next) data science role, based on thousands of reader questions.<br><strong>&#128073; <a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">Check out our </a></strong><em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">&#8220;Get A Data Science Job&#8221;</a></strong></em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide"> Course</a></strong><br></p></li><li><p><strong>Promote your organization/project/event to ~68,500 subscribers<br></strong>Sponsor this newsletter and reach a highly engaged data science audience (30&#8211;35% open rate).<br><strong>&#128073; Reply to this email to learn more</strong></p></li></ol><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 653]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-653</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-653</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 28 May 2026 19:41:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aks6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #653<br>May 28, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://fedemagnani.github.io/math/2026/04/08/the-quadratic-sandwich.html">The quadratic sandwich</a><br></strong>If you have ever tried to minimize a function with gradient descent, you probably noticed that some functions are a joy to optimize and others are a nightmare. The difference often boils down to two properties: strong convexity and L-smoothness. These two concepts define a &#8220;sandwich&#8221; of quadratic bounds around your function that tells you exactly how well-behaved it is. If the sandwich is tight, life is good. If one slice of bread is missing, things get ugly fast&#8230;In this post we&#8217;ll build up both concepts from scratch, see how they combine into the quadratic sandwich, understand what happens at the level of the Hessian&#8217;s eigenvalues, and pick up a neat trick to verify L-smoothness without ever computing an eigenvalue&#8230;</p></li></ul><ul><li><p><strong><a href="https://gudok.xyz/transpose/">What it takes to transpose a matrix</a></strong><br>In this article we are going to gradually build a sequence of progressively more efficient implementations of matrix transpose, with the most sophisticated implementation being up to x25 times faster than the naive one. During each step we will locate the bottleneck, figure out what has caused it, and think of a solution to overcome it. This article is intended to serve as an introduction to optimizing matrix algorithms for x86_64, presented from the perspective of a real-world problem&#8230;</p><p></p></li><li><p><strong><a href="https://www.johndcook.com/blog/2017/11/08/why-is-kullback-leibler-divergence-not-a-distance/">Why is Kullback-Leibler divergence not a distance?</a></strong><br>The Kullback-Leibler divergence between two probability distributions is a measure of how different the two distributions are. It is sometimes called a distance, but it&#8217;s not a distance in the usual sense because it&#8217;s not symmetric. At first this asymmetry may seem like a bug, but it&#8217;s a feature. We&#8217;ll explain why it&#8217;s useful to measure the difference between two probability distributions in an asymmetric way. The Kullback-Leibler divergence between two random variables X and Y is defined as&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:520345}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aks6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aks6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png 424w, https://substackcdn.com/image/fetch/$s_!aks6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png 848w, https://substackcdn.com/image/fetch/$s_!aks6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png 1272w, https://substackcdn.com/image/fetch/$s_!aks6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aks6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png" width="575" height="352" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:352,&quot;width&quot;:575,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:26143,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/199621460?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aks6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png 424w, https://substackcdn.com/image/fetch/$s_!aks6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png 848w, https://substackcdn.com/image/fetch/$s_!aks6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png 1272w, https://substackcdn.com/image/fetch/$s_!aks6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b83afc3-9536-455b-9f00-9275c0f64179_575x352.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://golfcoursewiki.substack.com/p/friday-pins-vs-sunday-pins-or-how">Friday Pins vs Sunday Pins or: How to Illustrate Something Completely Obvious</a></strong><br>In my previous article, I Spent the Last Month and a Half Building a Model that Visualizes Strategic Golf, I laid out the very basics of the golf model I built, the underlying reason that compelled me to work on it, and novel maps it could create. However, I barely scratched the surface of what this model can illustrate about golf course architecture. Here, I want to talk about how we can specifically look at one kind of dynamic architectural interest. That is, features of golf architecture that appear as we change the course setup. Specifically, I want to look at how different hole locations change the architectural interest for players, how we can show that, and what it looks like&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1tknjuv/what_ds_job_market_trends_are_you_seeing/">What DS job market trends are you seeing? [Reddit]</a></strong></p><p>I have 20 YOE, but I do a generic &#8220;data science&#8221; search on LinkedIn every 3 months to see how the job market is trending. Here are my latest observations. I would love to hear what others think.</p><ol><li><p>The number of AI postings is going down. ML and DE skills are back in fashion.</p></li><li><p>Salaries are down across the board.</p></li><li><p>Non-technical responsibility is up. I see &#8220;Data Scientist&#8221; roles being asked to create a roadmap and drive organizational change. That used to be the responsibility of the manager or maybe the lead.</p></li></ol><p>I haven&#8217;t applied for any of these jobs, so I don&#8217;t know what&#8217;s actually real. I wonder if Data Science is no longer the hot keyword and I should be searching for something else&#8230;</p><p></p></li><li><p><strong><a href="https://www.fharrell.com/talk/ai/">Thoughts About the Roles of AI for Statistics</a><br></strong>This talk covers what I&#8217;ve learned from using large language models in my work for the past two years. For statistical programming, success has come when I play the role of specification writer and comprehensive tester. For statistical methodology, AI has been successful serving me as a mathematical statistical assistant and a critic. Instead of avoiding AI we should embrace it, but we should always set a higher bar for the quality of our work as a result&#8230;<br></p></li><li><p><strong><a href="https://stats.stackexchange.com/questions/676087/evaluating-detection-classifier-algorithm-accuracy">Evaluating detection &amp; classifier algorithm accuracy</a><br></strong>Let&#8217;s say i have images with a mixture of normal cells and sick cells on each image. Humans can reliably distinguish normal cells from sick cells, however it takes a lot of time to mark up the images as there are hundreds of cells in one field of view. I have an algorithm that can also distinguish normal and sick cells. The outputs from both manual markup and my algorithm is 2 lists of (x, y) coordinates -- one list for sick cells, one for healthy. What are the best practices for comparing and reporting the accuracy of my algorithm against manual markup?&#8230;<br></p></li><li><p><strong><a href="https://www.statsignificant.com/p/do-most-tv-shows-stick-the-landing">Do Most TV Shows Stick the Landing?</a></strong></p><p>Four decades later, television has changed dramatically, reshaped by streaming and a clearer understanding of what makes for a satisfying conclusion. But has this institutional knowledge led to better endings? Have showrunners learned from the mistakes of St. Elsewhere, Game of Thrones, and other finale fiascos? So today, we&#8217;ll investigate whether omniscient showrunner Tommy Westphall has gotten any better at sticking the landing, how finale quality has changed in recent decades, and whether finality is simply a structural weakness of television itself&#8230;<br></p></li><li><p><strong><a href="https://redwallanalytics.com/posts/2023-02-22-nyed-data-explorer-shows-15-years-of-charter-school-success/">NYED Data Explorer Shows 15 Years of Charter School Success</a></strong><br>When I discovered 15 years of NYED assessment data, the interest to clean and free this data for others to discover in a Shiny app, was immediate. The opportunity to also feature Classical&#8217;s stand-out performance didn&#8217;t hurt my motivation, although this post and the app were built in my spare time, and do not represent the opinions of South Bronx Classical Charter Schools. Unlike many past Redwall posts, this one will not have code, and will be primarily to explain the data and show how to use the app&#8230;<code><br></code></p></li><li><p><strong><a href="https://statmodeling.stat.columbia.edu/2014/08/14/luck-vs-skill-poker/">Luck vs. skill in poker</a><br></strong>The thread of our recent discussion of quantifying luck vs. skill in sports turned to poker, motivating the present post&#8230;Can good poker players really &#8220;read&#8221; my cards and figure out what&#8217;s in my hand?&#8230;<br></p></li><li><p><strong><a href="https://remlapmot.github.io/post/2026/stan-compile-speedup/">Speeding up Stan model builds for R package developers</a><br></strong>My PhD student was interested in Bayesian methods and we put together an R package which included some Stan models. I was always frustrated by how slowly these compiled on our Windows machines&#8230;A few years later, when I got a MacBook Air I was shocked how much faster they compiled. On my Windows machine our mrbayes package takes 3 minutes 55 seconds to compile and install. On my M4 MacBook Air it takes 1 minute 16 seconds. The following tips show how to improve those timings&#8230;<br></p></li><li><p><strong><a href="https://ankitg.me/blog/2025/01/06/unfair-coins.html">How unfair is the coin?</a></strong></p><p>In February 2024, Reverie Labs, the startup I co-founded in 2017, was acquired by Ginkgo Bioworks. I&#8217;m now on leave from Ginkgo and I&#8217;ve joined Y Combinator as a Visiting Partner, giving me the chance to work with the next generation of companies. Especially in this new role, I&#8217;ve been thinking a bit about what worked, what didn&#8217;t work, and what lessons I can take forward&#8230;We had quite the journey &#8211; 6+ years of building at the intersection of AI and drug discovery. We began as a machine learning driven software company selling SaaS tools and consulting services to pharma companies, and at acquisition we were a pharmaceutical company, developing our own in-house pipeline of drug assets and advancing them rapidly using our machine learning technology&#8230;<br></p></li><li><p><strong><a href="https://jcarroll.com.au/2026/05/22/functions-over-idioms-rfuns/">Functions over Idioms - Writing R in Python with rfuns</a></strong><br>Sometimes a problem calls for a particular language to be used, and with that comes adjusting one&#8217;s brain to thinking in that language and using the appropriate idioms to leverage that language&#8217;s features&#8230;But what if I don&#8217;t want to?&#8230;The line between R and Python has been heavily blurred the last few years, particularly with {reticulate} (rstudio.github.io) enabling us to use Python within R code, RStudio rebranding as Posit (posit.co) and taking on a strong Python development effort, releasing Positron (posit.co) as a multi-language IDE, and Quarto (quarto.org) being a multi-language rethink of Rmarkdown&#8230;<br></p></li><li><p><strong><a href="https://aiweekender.substack.com/p/6-llm-prompting-techniques-for-data">LLM Prompting Techniques for Data Scientists and Engineers in 2026</a></strong></p><p>Six techniques matched to six failure modes, including inconsistent output formats, shallow reasoning, instruction drift, and more&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/selfimprovement/comments/1tjef6k/i_am_faking_my_way_through_a_data_analyst_role/">I am faking my way through a Data Analyst role with AI, how do I actually learn before I get caught? [Reddit]</a><br></strong>I graduated with a CS degree, but I spent my undergrad years grinding part-time jobs instead of actually studying. Now I am a Data Analyst at a small business, and the job is nothing like the theory I slept through in school. I am just winging it every day tbh. I rely heavily on openclaw for data scraping and acciowork to handle the processing and archiving. If these AI tools ever went down, I would be fired within an hour. I am terrified of being exposed as a fraud. Where do I even start fixing this? Should I grind python, or is mastering excel still the first step for survival?&#8230;<br></p></li><li><p><strong><a href="https://www.dougmacdowell.com/50-hours-to-draw-some-lines.html">50 Hours to Draw Some Lines</a><br></strong>"What are you working on these days?"<br>"Data visualizations." I told him.<br>"Ah, you using algorithms, machine learning, cloud computing, things like that?"<br>"No." I said. "I'm just trying to draw a line graph."&#8230;.What do I mean by drawing data by hand? I made this data visualization (data viz) about a coffee maker computer by hand, using rulers, pencils, ink, and a lettering kit. Along with my flubs, flukes, and acclimation with tools - it took me 50 hours to make. It&#8217;s statistically accurate, carefully crafted, and like Hackaday said &#8220;right out of a 1970&#8217;s college textbook&#8221;. It&#8217;s how professionals might visualize data before computers could do it for them&#8230;.</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1th87u4/are_there_any_small_quick_things_i_can_do/">Are there any small, quick things I can do everyday to keep my skills sharp? [Reddit]</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1tjbn57/after_5_years_in_data_science_im_starting_to/">After 5 years in data science, I&#8217;m starting to realize most &#8220;insights&#8221; we deliver are completely ignored. Is this normal? [Reddit]</a></strong></p></li><li><p><strong><a href="https://datascienceconfidential.github.io/r/predictive-models/2026/05/14/is-logistic-regression-regression.html">Is logistic regression regression?</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** Please take a look at last week's issue #652 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-652">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://sabr.org/lahman-database/">Lahman Baseball Database, created by SABR member Sean Lahman, contains complete major league batting and pitching statistics back to 1871</a></strong></p></li><li><p><strong><a href="https://mlbfranchiseanalysis.netlify.app/">A Data-Driven Survey of MLB Franchise Management</a></strong></p></li><li><p><strong><a href="https://jakubsobolewski.com/blog/bdd-shiny-when/">Behavior-Driven Development in R Shiny: Modeling User Behavior with When Steps</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/learnmath/comments/1t2mevm/i_dont_understand_standard_deviation/">I dont understand Standard Deviation [Reddit]</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1tmfjlw/good_practices_in_data_scripts/">Good practices in data scripts [Reddit]</a></strong></p></li></ul><p>.</p><div><hr></div><h2><strong>Whenever you're ready, 3 ways we can help:</strong><br></h2><ol><li><p><strong>Go deeper each week (paid subscription)</strong><br>Get 3 additional posts per week designed to help you:</p><ul><li><p>Statistics &#8594; understand the math behind ML</p></li><li><p>AI Agents &#8594; build with modern AI tools</p></li><li><p>Career &#8594; become more valuable at your job</p></li></ul><p><strong>&#128073; <a href="https://datascienceweekly.substack.com/subscribe">Upgrade for $10/month &#8212; cancel anytime</a><br></strong></p></li><li><p><strong>Looking to get a job?</strong><br>A practical guide to landing your first (or next) data science role, based on thousands of reader questions.<br><strong>&#128073; <a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">Check out our </a></strong><em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">&#8220;Get A Data Science Job&#8221;</a></strong></em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide"> Course</a></strong><br></p></li><li><p><strong>Promote your organization/project/event to ~68,500 subscribers<br></strong>Sponsor this newsletter and reach a highly engaged data science audience (30&#8211;35% open rate).<br><strong>&#128073; Reply to this email to learn more</strong></p></li></ol><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Who is going to save you?]]></title><description><![CDATA[On the importance of having a corporate sponsor and not putting all of your political eggs in one basket.]]></description><link>https://datascienceweekly.substack.com/p/who-is-going-to-save-you</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/who-is-going-to-save-you</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Fri, 22 May 2026 22:06:43 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" width="3000" height="2003" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2003,&quot;width&quot;:3000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;two women sits of padded chairs while using laptop computers&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="two women sits of padded chairs while using laptop computers" title="two women sits of padded chairs while using laptop computers" srcset="https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1573164574048-f968d7ee9f20?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Image Caption: <a href="https://unsplash.com/photos/two-women-sits-of-padded-chairs-while-using-laptop-computers-HocFQHhGjDE">Christina @ wocintechchat.com M</a></em></figcaption></figure></div>
      <p>
          <a href="https://datascienceweekly.substack.com/p/who-is-going-to-save-you">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 652]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-652</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-652</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 21 May 2026 22:52:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!I8ji!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3294ef46-cb03-42ea-b7d6-9f8e8b0f41f6_253x253.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #652<br>May 21, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://ccli.substack.com/p/whats-going-on-in-computational-neuroscience">What&#8217;s going on in computational neuroscience nowadays? (part 1)</a><br></strong>A month ago I came back from Cosyne, the annual Computational and Systems Neuroscience conference&#8230;The days are a haze of tutorials, talks, poster sessions, and workshops, usually appended by dinners and drinks past midnight&#8230;I seem to have a hard time writing one-off pieces, so I&#8217;m leaning into that and writing this as a series. I&#8217;ll only be able to write a very narrow perspective of Cosyne, of course, but most of the talks are on YouTube on the official channel if you want to see for yourself. (That also means this will be more personal thoughts than report)&#8230;</p></li></ul><ul><li><p><strong><a href="https://datascienceconfidential.github.io/r/predictive-models/2026/05/14/is-logistic-regression-regression.html">Is logistic regression regression?</a></strong><br>I came across a post recently by a machine learning engineer who made the bold claim that logistic regression is the worst name for an algorithm ever, or something along those lines&#8230;Many statisticians of the more old-school type seemed to disagree. This led me to think a bit more deeply about the subject. I&#8217;ve already written several posts on bad terminology in statistics (see confidence level, line of best fit, r squared) so I might have been expected to agree with the machine learning view, but in this case I agree with the statisticians, and I would like to explain why&#8230;</p><p></p></li><li><p><strong><a href="https://spawn-queue.acm.org/doi/pdf/10.1145/3778029">What Every Experimenter Must Know About Randomization</a></strong><br>Randomized controlled experiments offer gold-standard insight into cause and effect. The knowledge that informs our most important decisions. Unfortunately, randomization in such experiments is often botched. Randomization errors silently invalidate the interpretation of experimental results, turning a fruitful quest for knowledge into a waste of time and money, or, worse, a wellspring of misinformation. Fortunately, these fatal errors are easy to spot and fix. So whether you&#8217;re a webmaster using A/B testing to increase engagement, a medical researcher evaluating vaccines, a factory manager exploring productivity improvements, or a scientist seeking the laws that govern nature or human affairs, read on&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:516379}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://yihui.org/en/2026/05/testthat-to-testit/">Converting testthat Tests to testit</a></strong><br>Back in 2013, I wrote about testing R packages when I first released testit. Thirteen years later, I still believe that unit testing should be nothing more than &#8220;tell me if something unexpected happened.&#8221; Recently I converted a large testthat test suite to testit, and I thought I&#8217;d share a practical guide for anyone considering the same move&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1tjbn57/after_5_years_in_data_science_im_starting_to/">After 5 years in data science, I&#8217;m starting to realize most &#8220;insights&#8221; we deliver are completely ignored. Is this normal? [Reddit]</a></strong></p><p>I&#8217;ve been in data science roles (both analytics and ML) for about 5 years now across a couple of companies. Lately I&#8217;ve been feeling a bit burned out because I keep seeing the same pattern&#8230;We spend weeks cleaning data, building dashboards, running statistical analysis, or training models&#8230; and then the stakeholders either:</p><ul><li><p>Say &#8220;thanks&#8221; and never use it</p></li><li><p>Cherry-pick the numbers that support their existing opinion</p></li><li><p>Or just completely ignore the findings and go with gut feel anyway</p></li></ul><p>The worst part is when leadership asks for a &#8220;data-driven decision&#8221; but they&#8217;ve already decided what they want to do&#8230;Am I alone in this? Or is this just the reality of data science in most companies?&#8230;</p><p></p></li><li><p><strong><a href="https://vickiboykis.com/2026/05/18/tagging-my-blog-posts-with-bertopic-and-llms/">Tagging my blog posts with BERTopic and LLMs</a><br></strong>I recently added tags to my blog using BERTopic and a mix of LLMs. You can see the tags in the sidebar to the right (or in the footer on mobile). I&#8217;ve done this before in 2023, with GGUF Mistral using llama-cpp, but never finished the project. Now, because the models have been getting so good, and my project was small, relatively well-defined, and easy to evaluate, the project took me about 6-10 hours over a month, using BERTopic, Claude Code, and Pi with Deepseek&#8230;<br></p></li><li><p><strong><a href="https://ericmjl.github.io/blog/2026/5/20/what-data-science-is-actually-about-in-the-age-of-ai/">What data science is actually about in the age of AI</a><br></strong>I reflect on the evolving role of data scientists in the age of AI and LLMs. I argue that our core mission remains rigorous measurement, not full-stack development. While AI tools make building easier, the real value comes from defining and evaluating what truly matters. I share why measurement should be led by those closest to the problem and how data scientists can best contribute. Are we losing sight of what makes data science essential in the rush to build with AI?&#8230;<br></p></li><li><p><strong><a href="https://emma-x1.github.io/writing/transformer-from-scratch">Transformer From Scratch</a></strong></p><p>I&#8217;ve wanted to dive deeper into the fundamentals of AI for a while now - it feels a little bit magical, and a little bit wrong, to operate alongside AI without a strong understanding of how the underlying mechanisms work. Naturally, I had to write a transformer, and Neel Nanda&#8217;s &#8221;GPT-2 From Scratch&#8221; was my resource of choice&#8230;This post is meant to document my process of learning and to address some of the questions I was curious about when implementing the transformer for the first time. It includes an overview of transformer basics and some of my intuitions, followed by some of the points of interest (transformer secrets, if you will) and challenges I ran into&#8230;<br></p></li><li><p><strong><a href="https://simonwillison.net/2026/May/19/5-minute-llms/">The last six months in LLMs in five minutes</a></strong><br>I presented this lightning talk at PyCon US 2026, attempting to summarize the last six months of developments in LLMs in five minutes&#8230;.<code><br></code></p></li><li><p><strong><a href="https://remlapmot.github.io/post/2026/runiverse-tips/">Five tips for managing your R-universe</a><br></strong>rOpenSci&#8217;s R-universe system is an open source platform allowing users to create their own CRAN-like universe of R packages&#8230;This post gives five tips I have developed to help manage my R-universe&#8230;<br></p></li><li><p><strong><a href="https://blog.skypilot.co/research-driven-agents/">Research-Driven Agents: What Happens When Your Agent Reads Before It Codes</a><br></strong>Coding agents working from code alone generate shallow hypotheses. Adding a research phase &#8212; arxiv papers, competing forks, other backends &#8212; produced 5 kernel fusions that made llama.cpp CPU inference 15% faster.<br></p></li><li><p><strong><a href="https://people.duke.edu/~hpgavin/SystemID/References/Ribeiro-KalmanFilter-2004.pdf">Kalman and Extended Kalman Filters: Concept, Derivation and Properties</a></strong></p><p>This report presents and derives the Kalman filter and the Extended Kalman filter dynamics. The general filtering problem is formulated and it is shown that, under linearity and Gaussian conditions on the systems dynamics, the general filter particularizes to the Kalman filter. It is shown that the Kalman filter is a linear, discrete time, finite dimensional time-varying system that evaluates the state estimate that minimizes the mean-square error&#8230;<br></p></li><li><p><strong><a href="https://errorstatistics.com/2026/05/11/39531/">How not to turn power on its head</a></strong><br>In giving some informal remarks about power at a seminar a couple of weeks ago, I proposed that the tendency to turn the notion of power on its head might be avoided by imagining we need to define a test&#8217;s error probabilities in terms of its power alone. We can refer to the power against the null hypothesis, rather than alluding to a type 1 error probability, for example&#8230;What do I mean by turning power on its head? I mean, at least here, supposing that a test provides poor evidence of discrepancies that the test has low power to detect&#8230;<br></p></li><li><p><strong><a href="https://rworks.dev/posts/atlas-learn-sphere/">The Atlas-Learn Approach to the Manifold Hypothesis</a></strong></p><p>The 2025 paper by Robinett et al., &#8216;Atlas-based Manifold Representations for Interpretable Riemannian Machine Learning&#8217;, provides an algorithm for fitting a low dimensional manifold from a point cloud by means of a novel algorithm for approximating an atlas of charts. This post illustrates the Atlas-Learn method by reconstructing a sphere from a 3D point cloud of naive random samples and works through some checks on accuracy&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1th87u4/are_there_any_small_quick_things_i_can_do/">Are there any small, quick things I can do everyday to keep my skills sharp? [Reddit]</a><br></strong>I&#8217;m sure everyone knows about the dilemma of AI at this point. We want to work faster but our skills are atrophying yada yada&#8230;as a junior data scientist, I feel like I barely had any skills to begin with. Now with my company forcing us to use AI, I feel like I&#8217;m not learning much. Now I&#8217;ve been doing leetcode, but I just don&#8217;t think it&#8217;s that applicable to my real job. I don&#8217;t have the bandwidth outside of work to do a project yet, since my company is working us to the bone. What are some quick habits/tools/websites/apps you recommend to keep your skills sharp?..<br></p></li><li><p><strong><a href="https://www.counting-stuff.com/the-measurement-of-loudness/">The Measurement of Loudness</a><br></strong>I&#8217;m not a sound engineer of any sort, but I enjoy music and have been blessed with decent hearing acuity, so I tend to pay attention to noises going on around me. Now, I know what you&#8217;re thinking! Surely, a measurement nerd interested in sound would have bought a cheap SPL meter off the internet and this is what this post is about. And you&#8217;d be wrong! Hah! Because this post goes a bit further off the deep end because SPL and the dB(a) scales that we commonly associate with &#8220;sound volume&#8221; always confused me when I tried to understand them. Like with how measuring color is really difficult because it&#8217;s at the intersection of a physical measurement and human perception (see: How the heck does one measure color?), sound is just as messy because it&#8217;s again physical measurements (sound pressure) mediated by the human auditory system. To measure loudness, we&#8217;re going to have to go back a bit in history&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://idlemachines.co.uk/essays/softmax">Softmax, can you really derive the Jacobian? And should you care?</a></strong></p></li><li><p><strong><a href="https://kyunghyuncho.me/teaching-fundamentals-of-machine-learning/">Teaching &lt;Fundamentals of Machine Learning&gt;</a></strong></p></li><li><p><strong><a href="https://link.springer.com/article/10.1007/s42113-026-00271-1">Illusions of Understanding in the Sciences</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** Please take a look at last week's issue #651 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-651">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://xianblog.wordpress.com/2026/05/13/the-vexing-hausdorff-measure/">the vexing Hausdorff measure</a></strong></p></li><li><p><strong><a href="https://robjhyndman.com/publications/mvhts.html">Multivariate reconciliation for hierarchical time series</a></strong></p></li><li><p><strong><a href="https://possiblywrong.wordpress.com/2026/05/18/comments-on-what-every-experimenter-must-know-about-randomization/">Comments on: What Every Experimenter Must Know About Randomization</a></strong></p></li><li><p><strong><a href="https://www.rilldata.com/blog/introducing-metrics-sql-a-sql-based-semantic-layer-for-humans-and-agents">Introducing Metrics SQL: A SQL-based semantic layer for humans and agents</a></strong></p></li><li><p><strong><a href="https://ropensci.org/blog/2026/04/08/r-universe-bioc/">Collaborating between Bioconductor and R-universe on Development of Common Infrastructure</a></strong></p></li></ul><p>.</p><div><hr></div><h2><strong>Whenever you're ready, 3 ways we can help:</strong><br></h2><ol><li><p><strong>Go deeper each week (paid subscription)</strong><br>Get 3 additional posts per week designed to help you:</p><ul><li><p>Statistics &#8594; understand the math behind ML</p></li><li><p>AI Agents &#8594; build with modern AI tools</p></li><li><p>Career &#8594; become more valuable at your job</p></li></ul><p><strong>&#128073; <a href="https://datascienceweekly.substack.com/subscribe">Upgrade for $10/month &#8212; cancel anytime</a><br></strong></p></li><li><p><strong>Looking to get a job?</strong><br>A practical guide to landing your first (or next) data science role, based on thousands of reader questions.<br><strong>&#128073; <a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">Check out our </a></strong><em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">&#8220;Get A Data Science Job&#8221;</a></strong></em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide"> Course</a></strong><br></p></li><li><p><strong>Promote your organization/project/event to ~68,500 subscribers<br></strong>Sponsor this newsletter and reach a highly engaged data science audience (30&#8211;35% open rate).<br><strong>&#128073; Reply to this email to learn more</strong></p></li></ol><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Tools: Why AI Agents Need Them]]></title><description><![CDATA[Tools allow AI agents to interact with the world]]></description><link>https://datascienceweekly.substack.com/p/tools-why-ai-agents-need-them</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/tools-why-ai-agents-need-them</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Wed, 20 May 2026 23:03:55 GMT</pubDate><enclosure url="https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw"><img src="https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" width="3000" height="2436" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2436,&quot;width&quot;:3000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;clothes iron, hammer, axe, flashlight and pitcher on brown wooden table&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="clothes iron, hammer, axe, flashlight and pitcher on brown wooden table" title="clothes iron, hammer, axe, flashlight and pitcher on brown wooden table" srcset="https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/reserve/oIpwxeeSPy1cnwYpqJ1w_Dufer%20Collateral%20test.jpg?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image Source: <a href="https://unsplash.com/photos/clothes-iron-hammer-axe-flashlight-and-pitcher-on-brown-wooden-table-IClZBVw5W5A">Todd Quackenbush</a></figcaption></figure></div>
      <p>
          <a href="https://datascienceweekly.substack.com/p/tools-why-ai-agents-need-them">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Monday Statistics: Why Averages Can Mislead]]></title><description><![CDATA[When your data is skewed, the &#8220;average&#8221; may not represent reality]]></description><link>https://datascienceweekly.substack.com/p/monday-statistics-why-averages-can</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/monday-statistics-why-averages-can</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Mon, 18 May 2026 22:32:14 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" width="3000" height="2003" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2003,&quot;width&quot;:3000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;short-furred white cat&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="short-furred white cat" title="short-furred white cat" srcset="https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1549096454-1b8ba2ef8f1c?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Image Source: <a href="https://unsplash.com/photos/short-furred-white-cat-diLm-niNsJk">Gaelle Marcel</a></em></figcaption></figure></div>
      <p>
          <a href="https://datascienceweekly.substack.com/p/monday-statistics-why-averages-can">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Figure Out Who Your Competition Is]]></title><description><![CDATA[If you applied for a job next week, you&#8217;d be competing with a set of candidates. You should find out who they are and what they are doing.]]></description><link>https://datascienceweekly.substack.com/p/figure-out-who-your-competition-is</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/figure-out-who-your-competition-is</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Fri, 15 May 2026 12:35:20 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" width="3000" height="2000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2000,&quot;width&quot;:3000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;person in red sweater holding babys hand&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="person in red sweater holding babys hand" title="person in red sweater holding babys hand" srcset="https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1582213782179-e0d53f98f2ca?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Image Source: <a href="https://unsplash.com/photos/person-in-red-sweater-holding-babys-hand-Zyx1bK9mqmA">Hannah Busing</a></em></figcaption></figure></div>
      <p>
          <a href="https://datascienceweekly.substack.com/p/figure-out-who-your-competition-is">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 651]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-651</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-651</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 14 May 2026 20:23:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!obNH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #651<br>May 14, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://www.jackhogan.me/blog/marco-polo">Marco Polo: Finding a friend with only distance and motion.</a><br></strong>You walk into a cafe, looking for your friend. Seems like an easy task, until you see it&#8217;s so packed that you can&#8217;t see through the crowd at all, and everyone&#8217;s talking so loud that you can barely hear anything. The only things you know are your movements, and how far you are from your friend (through the special psychic bond you two share). How will you find each other?...I wanted to solve the exact same problem, but with devices instead of people (so no psychic connection for me), existing in a space of hundreds of other devices. Working the problem taught me a lot of really interesting science relating to robotics and state estimation, and I wrote this post so you can learn, too&#8230;</p></li></ul><ul><li><p><strong><a href="https://burrito.bio/essays/biology-is-a-burrito">Biology is a Burrito</a></strong><br>A bacterium&#8217;s genome, pulled into a straight thread, is nearly 1,000 times longer than the cell from which it came. If you placed one E. coli into a gallon-sized jug with some nutrients and waited a few hours, the genomes of its descendants, placed end-to-end, would reach to the moon and back...several times&#8230;.The truth is that biochemistry textbooks often depict cells as spacious places, where molecules float in secluded harmony. &#8220;But a cell looks more like a burrito,&#8221; says Michael Elowitz, a biologist at Caltech. All the biochemicals are pushed together, bumping into each other&#8230;</p><p></p></li><li><p><strong><a href="https://idlemachines.co.uk/essays/softmax">Softmax, can you really derive the Jacobian? And should you care?</a></strong><br>Multiclass output? Softmax. Normalising probabilities? Softmax. Attention weights? Softmax. Partition function? You guessed it, Softmax. This function comes up everywhere, but how often have you really thought about what&#8217;s going on inside?&#8230;What does softmax actually do to your distribution?&#8230;The softmax function is deceptively simple&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><p></p><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!obNH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!obNH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png 424w, https://substackcdn.com/image/fetch/$s_!obNH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png 848w, https://substackcdn.com/image/fetch/$s_!obNH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png 1272w, https://substackcdn.com/image/fetch/$s_!obNH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!obNH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png" width="600" height="353.8732394366197" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:670,&quot;width&quot;:1136,&quot;resizeWidth&quot;:600,&quot;bytes&quot;:68648,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/197745894?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!obNH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png 424w, https://substackcdn.com/image/fetch/$s_!obNH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png 848w, https://substackcdn.com/image/fetch/$s_!obNH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png 1272w, https://substackcdn.com/image/fetch/$s_!obNH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cde0e3c-58c5-49de-8783-47919460d244_1136x670.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://link.springer.com/article/10.1007/s42113-026-00271-1">Illusions of Understanding in the Sciences</a></strong><br>The first part of this essay supports the case for the universality of partial and incomplete levels of understanding by showing the difficulty of reaching a deep level of understanding for even a simple analysis and model that most scientists use and believe they understand: linear regression. The second part highlights some implications of the existence of many levels of understanding and explanation, and their use by scientists for design, testing, analysis, and theory development. It discusses the way that deduction and induction depend on the levels of understanding and the implications of the illusion that a scientist&#8217;s understanding is deep&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/MachineLearning/comments/1t72u1r/people_interested_in_continual_learning_researchr/">People Interested in Continual Learning Research [Reddit]</a></strong></p><p>Recently, I&#8217;ve become fascinated by Continual Learning, especially the idea of AI systems that can continuously adapt and improve from experience rather than staying static after training. I&#8217;m a student just starting my journey in CL research and would love to connect with people exploring similar ideas. Whether you&#8217;re a student, researcher, or just curious about the field, feel free to DM me. Would also love paper recommendations and interesting research directions&#8230;</p><p></p></li><li><p><strong><a href="https://www.youtube.com/watch?v=eSZ9FB5Dqnk">You Should Probably Map That: Introduction to Geospatial Analysis in R</a><br></strong>Anjile An, Weill Cornell Medical College&#8230;Data comes in many forms, but the spatial components can be overlooked. You&#8217;ll get some history of mapping and learn the basic building blocks of spatial data.  The demo portion will go over how to use the trusty {ggplot2} as well as some new tools like {sf} and {tmap} to plot different types of spatial data&#8230;<br></p></li><li><p><strong><a href="https://kiro.dev/blog/deep-spec-analysis/">Requirements analysis: catching requirement bugs before they become code</a><br></strong>Every experienced engineer has a story where a feature shipped, worked on the happy path, and then quietly did the wrong thing on some edge case no one had thought about. Trace the bug back far enough and it rarely ends at the code; it ends at a sentence in a requirement document that meant one thing to the person who wrote it and something else to the person who implemented it&#8230;<br></p></li><li><p><strong><a href="https://kyunghyuncho.me/teaching-fundamentals-of-machine-learning/">Teaching &lt;Fundamentals of Machine Learning&gt;</a></strong></p><p>This past spring, i taught &lt;Fundamentals of Machine Learning&gt; for computer science seniors (with some juniors as well as seniors from other majors, including data science and economics) at NYU. last time i taught this course, the course was titled &lt;Introduction to Machine Learning&gt;, it was pre-ChatGPT and it was pre-pandemic; in fact, i was teaching this course in the spring of 2020, and the whole university, city and world went into its first lock down mid-way. in other words, i taught this course in the old world, and i was asked to teach this course in this brave new world&#8230;<br></p></li><li><p><strong><a href="https://robotchinwag.com/posts/jensen-shannon-divergence-visualisation/">Interactive Jensen&#8211;Shannon Divergence Visualisation</a></strong><br>An interactive visualisation of Jensen&#8211;Shannon divergence - the symmetric, always-finite cousin of KL. Shape two distributions and watch JSD, its ceiling of one bit, and the per-point contribution respond in real time&#8230;<code><br></code></p></li><li><p><strong><a href="https://puntofisso.net/eurovision/">70 years of love, empowerment, and freedom. Here&#8217;s a look at Eurovision by its lyrics.</a><br></strong>Famous for its upbeat rhythms, flamboyant performances, and political voting patterns, the competition has another dimension that moves with the zeitgeist: its song lyrics&#8230;<br></p></li><li><p><strong><a href="https://margaretstorey.com/blog/2026/02/18/cognitive-debt-revisited/">What I&#8217;m Hearing About Cognitive Debt (So Far)</a><br></strong>A week ago, I wrote about how Generative and Agentic AI may be amplifying what I&#8217;ve been calling cognitive debt: the accumulated gap between a system&#8217;s evolving structure and a team&#8217;s shared understanding of how and why that system works and can be changed over time. The post sparked thoughtful discussion across different communities. Rather than respond thread by thread, I want to synthesize what I&#8217;m hearing and connect it to other reflections I&#8217;ve been reading. I will likely update this as the conversation evolves&#8230;<br></p></li><li><p><strong><a href="https://ivanpleshkov.dev/blog/polynomial-autoencoder/">Polynomial autoencoder</a></strong></p><p>The most direct way to compress an embedding (other than quantization) is to fit PCA on the corpus and keep the top-d eigenvectors. It works, but PCA is a linear projection, and neural-network embeddings on the sphere are structurally nonlinear &#8212; the well-known cone effect in transformers. Some of the variance lives in a nonlinear tail that a linear decoder can&#8217;t reach&#8230;This post is about a closed-form way to add a quadratic decoder on top of PCA, to capture part of that nonlinear tail. The encoder stays as plain PCA. The decoder is a degree-2 polynomial lift plus Ridge OLS (ordinary linear regression with L2 regularization), also closed-form. No SGD, no epochs, no hyperparameter search. One np.linalg.solve over corpus statistics&#8230;<br></p></li><li><p><strong><a href="https://brrrviz.com/">Visualize the Brrr - Learn GPU programming</a></strong><br><strong>Who This Is For: </strong>This is for anyone interested in GPU programming or performance engineering, whether you&#8217;re writing CUDA, HIP, Triton, cuTile, Gluon, Helion, JAX, or nothing at all. If you&#8217;ve ever struggled to conceptualize parallelism, memory coalescing, or tiling, these visualizations are meant to make those ideas concrete. No prior GPU experience required.</p><p><strong>What You&#8217;ll Learn: </strong>The lessons walk through fundamental concepts of GPU programming, parallelism, memory hierarchy, and more. Each chapter pairs a short explanation with an interactive visualization you can poke at&#8230;<br></p></li><li><p><strong><a href="https://github.com/RussellSB/pytrendy">PyTrendy</a></strong></p><p>PyTrendy is a robust solution for identifying and analyzing trends in time series. Unlike other trend detection packages, it is robust to noisy &amp; flat segments, and handles for gradual &amp; abrupt trend cases with a high precision. It aims to be the best package for trend detection in python&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1sw6b6w/standardization_vs_log_transform/">Standardization vs Log transform? [Reddit]</a><br></strong>I have been trying to understand the use cases of both of these and I am really confused. I know log transform fixes the features and makes their distribution normal and standardization on the other hand only fixes the scale of the feature by keeping the distribution the same. Are these things which I use one after the other ? Or just simply use one depending on the case (which I also don&#8217;t understand when) ?&#8230;<br></p></li><li><p><strong><a href="https://www.seascapemodels.org/posts/2026-03-28-agentic-AI-ecological-modelling/">AI agents can create convincing ecological models, but you still need to know what you&#8217;re doing</a><br></strong>Agentic AI tools like Claude Code can write and run code, fix its own errors, and produce a formatted report with figures. I wanted to know whether that translates into reliable ecological modelling, so we ran a test: three fisheries tasks, four AI models, ten independent runs each, scored against a rubric. The results are published in <em>Fish and Fisheries</em>. We found agents can be genuinely useful, but only if you know how to use them well and only if you know enough about the analysis to catch what they miss&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1t2nasr/a_decade_of_being_an_average_data_scientist_my/">A decade of being an average Data Scientist! My personal experience [Reddit].</a></strong></p></li><li><p><strong><a href="https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs">Notes from inside China&#8217;s AI labs: Lessons from my trip to talk to most of the leading AI labs in China.</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1t19v2s/ds_market_is_kind_of_insane_right_now/">DS market is kind of insane right now [Reddit]</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** Please take a look at last week's issue #650 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-650">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://arxiv.org/abs/1912.10642">Notes on Category Theory with examples from basic mathematics</a></strong></p></li><li><p><strong><a href="https://planetscale.com/learn/courses/mysql-for-developers">MySQL for Developers</a></strong></p></li><li><p><strong><a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/">Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity</a></strong></p></li><li><p><strong><a href="https://www.seascapemodels.org/posts/2026-04-23-clear-plots-ggplot-for-mobile/">Design your plots (ggplot) for mobile</a></strong></p></li><li><p><strong><a href="https://rtichoke.netlify.app/posts/assessing-score-reliability.html">Assessing Credit Score Prediction Reliability Using Bootstrap Resampling</a></strong></p></li><li><p><strong><a href="https://appliedscientific.ai/research/scientific-ai-literature-agent-nvidia-nemotron-nano-omni">The Figure Problem in Scientific AI: Building a Multimodal Literature Agent for Biology, Powered by NVIDIA Nemotron 3 Nano Omni</a></strong></p></li><li><p><strong><a href="https://mfatihtuzen.github.io/posts/2026-04-24_quarto_blog_github/">Publishing a Quarto Blog: What I Learned Moving from Netlify to GitHub Pages</a></strong></p></li></ul><p>.</p><div><hr></div><h2><strong>Whenever you're ready, 3 ways we can help:</strong><br></h2><ol><li><p><strong>Go deeper each week (paid subscription)</strong><br>Get 3 additional posts per week designed to help you:</p><ul><li><p>Statistics &#8594; understand the math behind ML</p></li><li><p>AI Agents &#8594; build with modern AI tools</p></li><li><p>Career &#8594; become more valuable at your job</p></li></ul><p><strong>&#128073; <a href="https://datascienceweekly.substack.com/subscribe">Upgrade for $10/month &#8212; cancel anytime</a><br></strong></p></li><li><p><strong>Looking to get a job?</strong><br>A practical guide to landing your first (or next) data science role, based on thousands of reader questions.<br><strong>&#128073; <a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">Check out our </a></strong><em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">&#8220;Get A Data Science Job&#8221;</a></strong></em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide"> Course</a></strong><br></p></li><li><p><strong>Promote your organization/project/event to ~68,500 subscribers<br></strong>Sponsor this newsletter and reach a highly engaged data science audience (30&#8211;35% open rate).<br><strong>&#128073; Reply to this email to learn more</strong></p></li></ol><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI Agents for Data Scientists: Automations vs Agents]]></title><description><![CDATA[Why boring workflows with one LLM call often outperform &#8220;fully autonomous&#8221; systems]]></description><link>https://datascienceweekly.substack.com/p/ai-agents-for-data-scientists-automations</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/ai-agents-for-data-scientists-automations</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Wed, 13 May 2026 14:30:10 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" width="3000" height="1995" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1995,&quot;width&quot;:3000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;red and green plastic toy&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="red and green plastic toy" title="red and green plastic toy" srcset="https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1607601191544-fd61c99dd3c9?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Image Source: <a href="https://unsplash.com/photos/red-and-green-plastic-toy-jS_T9wWFTyE">Brian Kamau</a></em></figcaption></figure></div>
      <p>
          <a href="https://datascienceweekly.substack.com/p/ai-agents-for-data-scientists-automations">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Monday Statistics: Skewed Data and Transformations]]></title><description><![CDATA[Why your data looks &#8220;lopsided,&#8221; and what to do about it]]></description><link>https://datascienceweekly.substack.com/p/monday-statistics-skewed-data-and</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/monday-statistics-skewed-data-and</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Mon, 11 May 2026 21:39:22 GMT</pubDate><enclosure url="https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw"><img src="https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D" width="3000" height="2000" data-attrs="{&quot;src&quot;:&quot;https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2000,&quot;width&quot;:3000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Modern building with glass and red brick facade&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Modern building with glass and red brick facade" title="Modern building with glass and red brick facade" srcset="https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 424w, https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 848w, https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1272w, https://images.unsplash.com/photo-1762889577454-6c29b20b0e9b?fm=jpg&amp;q=60&amp;w=3000&amp;auto=format&amp;fit=crop&amp;ixlib=rb-4.1.0&amp;ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Image Source: <a href="https://unsplash.com/photos/modern-building-with-glass-and-red-brick-facade-awLg1eJFhaU">Pix Tresa</a></em></figcaption></figure></div>
      <p>
          <a href="https://datascienceweekly.substack.com/p/monday-statistics-skewed-data-and">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 650]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-650</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-650</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 07 May 2026 20:50:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gZvZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #650<br>May 07, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://www.interconnects.ai/p/notes-from-inside-chinas-ai-labs">Notes from inside China&#8217;s AI labs</a><br></strong>Lessons from my trip to talk to most of the leading AI labs in China&#8230;</p></li></ul><ul><li><p><strong><a href="https://www.kenkoonwong.com/blog/survival/">Learning &amp; Exploring Survival Analysis Part 1 - A Note To Myself</a></strong><br>A note to myself on survival analysis &#8212; KM curves, log-rank tests &amp; Cox models &#129518; If I wrote it the way I understood it, maybe I&#8217;ll actually remember it &#129310;&#8230;</p><p></p></li><li><p><strong><a href="https://www.youtube.com/watch?v=7OJlS2NW00g">Two Cartography Expert Review MOVIE MAPS</a></strong><br>John Nelson and Peter Attwood review maps in films. What they get right, what they might get wrong, and what we think of them as cartography nerds. Includes thoughts on: Indiana Jones, The Lord of the Rings, Avatar, War Games, Prometheus, Harry Potter, The Muppets, Pirates of the Caribbean, Game of Thrones, Moonrise Kingdom, The Goonies, and The Emperor&#8217;s New Groove&#8230;</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:508882}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gZvZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gZvZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png 424w, https://substackcdn.com/image/fetch/$s_!gZvZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png 848w, https://substackcdn.com/image/fetch/$s_!gZvZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png 1272w, https://substackcdn.com/image/fetch/$s_!gZvZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gZvZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png" width="600" height="326.8041237113402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:634,&quot;width&quot;:1164,&quot;resizeWidth&quot;:600,&quot;bytes&quot;:82342,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/196824414?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gZvZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png 424w, https://substackcdn.com/image/fetch/$s_!gZvZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png 848w, https://substackcdn.com/image/fetch/$s_!gZvZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png 1272w, https://substackcdn.com/image/fetch/$s_!gZvZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a305f6c-c067-4773-b025-2fb7676fde03_1164x634.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://kieranhealy.org/blog/archives/2026/05/02/bad-weather-and-the-subway/">Bad Weather and the Subway</a></strong><br>I&#8217;ve been looking at hourly ridership data from the New York City Subway. Last time we learned that people go to work in the morning and come home in the evening, for example. (All together now: &#8220;Only in New York, baby!&#8221;) Today, we&#8217;ll learn that bad weather makes people stay at home. Except, sometimes it doesn&#8217;t&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1t2nasr/a_decade_of_being_an_average_data_scientist_my/">A decade of being an average Data Scientist! My personal experience. [Reddit]</a></strong></p><p>Hello! I know there&#8217;s people here with PhDs, working in FAANG, on top of the newest tech, and are absolutely brilliant Data Scientists. I&#8217;m not one of them&#8230;I just wanted to share my positive experience from someone who is painfully average lol!! I wanted to show people, especially new grads and/or people pivoting into the field, that you don&#8217;t have to be the smartest person in the room to get hired. You need to drill into the solid foundations and a have a drive to make change/bring value to a company&#8230;</p><p></p></li><li><p><strong><a href="https://www.goodfire.ai/research/the-world-inside-neural-networks#">The World Inside Neural Networks</a><br></strong>How neural geometry will unlock understanding and control of AI&#8230;If we understand how a model carves up and represents the world conceptually (i.e., its ontology), this will unlock a far deeper understanding of both its algorithms that operate over that ontology and the intelligent behaviors produced by those algorithms. We thus need new methods to gain that understanding; this series of posts details our early efforts to develop them, building upon &#8211; and alongside &#8211; numerous related efforts from others&#8230;<br></p></li><li><p><strong><a href="https://pawankjha.substack.com/p/quantum-machine-learning-the-pragmatic">Quantum Machine Learning: The Pragmatic Guide for classical ML Engineers</a><br></strong>Part 1 of the &#8220;Quantum ML for Engineers&#8221; series: From Transformers and GPUs to QPUs and Hybrid Intelligence&#8230;<br></p></li><li><p><strong><a href="https://opensource.posit.co/blog/2026-05-07_opentelemetry/">Bringing OpenTelemetry to R in production</a></strong><br>Posit has instrumented Shiny, plumber2, mirai, httr2, ellmer, knitr, testthat, and DBI with OpenTelemetry, and created tools for you to instrument your own packages, bringing production-grade observability to R&#8230;<br></p></li><li><p><strong><a href="https://ds100.org/sp26/">Data 100: Principles and Techniques of Data Science UC Berkeley, Spring 2026</a></strong><br>This intermediate level class bridges between Data 8 and upper division computer science and statistics courses as well as methods courses in other fields. In this class, we explore key areas of data science including question formulation, data collection and cleaning, visualization, statistical inference, predictive modeling, and decision making.&#8203; Through a strong emphasis on data centric computing, quantitative critical thinking, and exploratory data analysis, this class covers key principles and techniques of data science&#8230;<code><br></code></p></li><li><p><strong><a href="https://nejsds.nestat.org/journal/NEJSDS/article/114/info">Inverse Probability Weighting: From Survey Sampling to Evidence Estimation</a><br></strong>We consider the class of inverse probability weight (IPW) estimators, including the popular Horvitz&#8211;Thompson and H&#225;jek estimators used routinely in survey sampling, causal inference and for Bayesian computation. We focus on the &#8216;weak paradoxes&#8217; for these estimators due to two counterexamples by Basu (1988) and Wasserman (2004) and investigate the two natural Bayesian answers to this problem: one based on binning and smoothing: a &#8216;Bayesian sieve&#8217; and the other based on a conjugate hierarchical model that allows borrowing information via exchangeability&#8230;<br></p></li><li><p><strong><a href="https://blog.ephorie.de/the-magic-of-in-context-learning-icl-when-your-model-already-knows-your-data">The Magic of In-Context Learning (ICL): When Your Model Already Knows Your Data</a></strong><br>As an experienced data scientist, you have seen thousands of datasets in your career. When confronted with new data, your natural neural network (a.k.a. brain) simply draws on this vast library of past mathematical shapes and immediately recognizes the pattern. But what if an artificial neural network could do exactly the same thing? What if it could predict your data without actually being trained on it?&#8230;Welcome to the mind-bending world of <em>In-Context Learning (ICL)</em> for tabular data, brought to R via the incredible new <code>TabPFN</code> package (on CRAN)&#8230;<br></p></li><li><p><strong><a href="https://www.johndcook.com/blog/2026/04/30/derivative-of-relu/">Three ways to differentiate ReLU</a></strong></p><p>When a function is not differentiable in the classical sense there are multiple ways to compute a generalized derivative. This post will look at three generalizations of the classical derivative, each applied to the ReLU (rectified linear unit) function. The ReLU function is a commonly used activation function for neural networks. It&#8217;s also called the ramp function for obvious reasons&#8230;<br></p></li><li><p><strong><a href="https://thierrymoudiki.github.io//blog/2026/05/02/r/rvflnet">You Don&#8217;t Need to Learn All the Weights on tabular data: The Case for rvflnet (a nonlinear expressive glmnet) on regression, classification and survival analysis</a></strong><br>Random Vector Functional Link (RVFL) networks offer a simple yet powerful alternative to traditional neural networks for tabular data. Instead of learning hidden layers through backpropagation, RVFL generates them randomly (or not, if using a deterministic sequence of quasi-random numbers) and focuses all learning effort on a final, regularized linear model&#8230;<br></p></li><li><p><strong><a href="https://jcarroll.com.au/2026/05/04/comparing-r-s-targets-and-dbt-for-data-engineering/">Comparing R&#8217;s {targets} and dbt for Data Engineering</a></strong></p><p>I&#8217;m getting more and more into data engineering these days and having used R for a long time, I&#8217;m seeing a lot of problems that look nail-shaped to my R-shaped hammer. The available tools to solve those problems exist for (presumably) very good reasons, so I wanted to take some time to dig into how to use them and compare their workflows to what I would otherwise naively do in R&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1t19v2s/ds_market_is_kind_of_insane_right_now/">DS market is kind of insane right now [Reddit]</a><br></strong>So here&#8217;s the story: another team in my company opened an associate-level DS role last week, we got 300+ applications, and somehow there were 30+ senior-level guys applying for it. Not fake senior either. Like actually senior all with 10+ yoe&#8230;.Curious that are other people &amp; teams seeing the same thing, or is this just a weird sample on our side?&#8230;<br></p></li><li><p><strong><a href="https://publicdomainreview.org/collection/visualizing-history-the-polish-system/">Visualizing History: The Polish System</a><br></strong>For the Polish educator Antoni Ja&#380;wi&#324;ski, history was best represented by an abstract grid &#8212; or at least it was for the purposes of remembering it. The so-called &#8220;Polish System&#8221; originated in the 1820s and was later brought to public attention in the 1830s and 1840s by General J&#243;zef Bem, a military engineer with a penchant for mnemonics&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1srp178/warning_dont_get_gptbrained/">Warning: Don&#8217;t get GPT-brained [Reddit]</a></strong></p></li><li><p><strong><a href="https://www.youtube.com/watch?v=sxX8BMscce0">Principles for Autonomous System Design: OpenClaw Deep Dive</a></strong></p></li><li><p><strong><a href="https://perthirtysix.com/how-the-heck-does-shazam-work">How The Heck Does Shazam Work?</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** Please take a look at last week's issue #649 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-649">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://www.allendowney.com/blog/2026/05/01/planning-for-your-midlife-crisis/">Planning for your midlife crisis (or Counterfactual Analysis with Bayesian Models: What Drives the Life Expectancy Gap?)</a></strong></p></li><li><p><strong><a href="https://blog.isquaredsoftware.com/2026/05/ai-thoughts-part-1-fears-opinions-journey/">My Thoughts on AI, Part 1: Fears, Opinions, and Mental Journey</a></strong></p></li><li><p><strong><a href="https://yihui.org/en/2026/05/ai-reflections/">Reflections on AI-assisted Programming</a></strong></p></li><li><p><strong><a href="https://mindfulmodeler.substack.com/p/time-series-forecasting-with-tabular">Time series forecasting with tabular foundation models</a></strong></p></li><li><p><strong><a href="https://www.statsignificant.com/p/should-you-trust-the-netflix-top">Should You Trust the Netflix Top 10? A Statistical Analysis</a></strong></p></li></ul><p>.</p><div><hr></div><h2><strong>Whenever you're ready, 3 ways we can help:</strong><br></h2><ol><li><p><strong>Go deeper each week (paid subscription)</strong><br>Get 3 additional posts per week designed to help you:</p><ul><li><p>Statistics &#8594; understand the math behind ML</p></li><li><p>AI Agents &#8594; build with modern AI tools</p></li><li><p>Career &#8594; become more valuable at your job</p></li></ul><p><strong>&#128073; <a href="https://datascienceweekly.substack.com/subscribe">Upgrade for $10/month &#8212; cancel anytime</a><br></strong></p></li><li><p><strong>Looking to get a job?</strong><br>A practical guide to landing your first (or next) data science role, based on thousands of reader questions.<br><strong>&#128073; <a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">Check out our </a></strong><em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">&#8220;Get A Data Science Job&#8221;</a></strong></em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide"> Course</a></strong><br></p></li><li><p><strong>Promote your organization/project/event to ~68,500 subscribers<br></strong>Sponsor this newsletter and reach a highly engaged data science audience (30&#8211;35% open rate).<br><strong>&#128073; Reply to this email to learn more</strong></p></li></ol><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Data Science Weekly - Issue 649]]></title><description><![CDATA[Curated news, articles and jobs related to Data Science, AI, & Machine Learning]]></description><link>https://datascienceweekly.substack.com/p/data-science-weekly-issue-649</link><guid isPermaLink="false">https://datascienceweekly.substack.com/p/data-science-weekly-issue-649</guid><dc:creator><![CDATA[Data Science Weekly]]></dc:creator><pubDate>Thu, 30 Apr 2026 21:37:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nTV4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!byfl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png" width="319" height="253" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17becea5-db12-4465-be92-858de78b9137_319x253.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:253,&quot;width&quot;:319,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Data Science Weekly&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Data Science Weekly" title="Data Science Weekly" srcset="https://substackcdn.com/image/fetch/$s_!byfl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 424w, https://substackcdn.com/image/fetch/$s_!byfl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 848w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1272w, https://substackcdn.com/image/fetch/$s_!byfl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17becea5-db12-4465-be92-858de78b9137_319x253.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2><strong>Issue #649<br>April 30, 2026<br></strong></h2><div><hr></div><p>Hello!</p><p><strong>Once a week, we write this email to share the links we thought were worth sharing in the Data Science, ML, AI, Data Visualization, and ML/Data Engineering worlds.</strong></p><div><hr></div><p><em><strong>And now&#8230;let&#8217;s dive into some interesting links from this week.</strong></em></p><div><hr></div><h2><strong>Editor's Picks<br></strong></h2><ul><li><p><strong><a href="https://arxiv.org/abs/2604.26268">The Difference Between &#8220;Replicable&#8221; and &#8220;Not replicable&#8221; is not Itself Scientifically Replicable</a><br></strong>Replication studies estimate the replicability rate of scientific results by aggregating binary verdicts of experiments. Exact replications are rarely attainable, so most replication sequences are non-exact. Experiments differ in ways that matter and do not share a single data-generating process. We formalize two statistical interpretations of non-exactness. In a shared latent rate (benchmark) model, experiments are exchangeable and depend on a common random replicability rate. In a conditionally independent rates (operational) model, each experiment has its own replicability rate drawn from a population distribution&#8230;&#8230;&#8230;The replication crisis, if there is one, cannot be established by the methods used to declare it&#8230;.</p></li></ul><ul><li><p><strong><a href="https://lemire.me/blog/2026/04/27/you-can-beat-the-binary-search/">You can beat the binary search</a></strong><br>Binary search is a classic algorithm that efficiently locates a target value in a sorted array by repeatedly dividing the search interval in half&#8230;In C++, this is implemented by the <code>std::binary_search</code> function, which returns a boolean indicating whether the value is present&#8230;The popular Roaring Bitmap format uses arrays of 16-bit integers of size ranging from 1 to 4096. We sometimes have to check whether a value is present. We use a binary search&#8230;I wanted a faster approach. I had two insights&#8230;Virtually all processors today have data parallel instructions (sometimes called SIMD) that can check several values at once&#8230;The binary search checks one value at a time. However, recent processors can load and check more than one value at once&#8230;Thus, I created something I call the SIMD Quad algorithm. It is an efficient search algorithm for sorted arrays of 16-bit unsigned integers, combining a quaternary interpolation search with SIMD (Single Instruction, Multiple Data)&#8230;</p><p></p></li><li><p><strong><a href="https://www.youtube.com/watch?v=sxX8BMscce0">Principles for Autonomous System Design: OpenClaw Deep Dive</a></strong><br>In this talk, Alex Krentsel (UC Berkeley, NetSys Lab / Google Research) does a deep-dive into OpenClaw &#8212; a fully open-source autonomous AI agent system &#8212; and uses it as a lens to explore the emerging design principles behind truly autonomous agents. We&#8217;re in Phase 3 of the AI evolution: LLM + tool-use + dynamic tool discovery. The agents that exist today aren&#8217;t chatbots. They read your email, write code, schedule work, remember context across sessions, and spawn other agents. This talk breaks down exactly how that works&#8230;.</p></li></ul><div><hr></div><h1><strong>What&#8217;s on your mind</strong></h1><h2>This Week&#8217;s Poll:</h2><div class="poll-embed" data-attrs="{&quot;id&quot;:504914}" data-component-name="PollToDOM"></div><p>.</p><h2>Last Week&#8217;s Poll:</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nTV4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nTV4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png 424w, https://substackcdn.com/image/fetch/$s_!nTV4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png 848w, https://substackcdn.com/image/fetch/$s_!nTV4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png 1272w, https://substackcdn.com/image/fetch/$s_!nTV4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nTV4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png" width="590" height="390.9138840070299" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:754,&quot;width&quot;:1138,&quot;resizeWidth&quot;:590,&quot;bytes&quot;:99059,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://datascienceweekly.substack.com/i/196043785?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nTV4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png 424w, https://substackcdn.com/image/fetch/$s_!nTV4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png 848w, https://substackcdn.com/image/fetch/$s_!nTV4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png 1272w, https://substackcdn.com/image/fetch/$s_!nTV4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93db8bd7-89d2-41e1-9018-6b1f622271c9_1138x754.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><div><hr></div><h2>Data Science Articles &amp; Videos</h2><p></p><ul><li><p><strong><a href="https://www.mdpi.com/1424-8220/22/5/1885">Linear Regression vs. Deep Learning: A Simple Yet Effective Baseline for Human Body Measurement</a></strong><br>We propose a linear regression model for the estimation of human body measurements. The input to the model only consists of the information that a person can self-estimate, such as height and weight. We evaluate our model against the state-of-the-art approaches for body measurement from point clouds and images, demonstrate the comparable performance with the best methods, and even outperform several deep learning models on public datasets. The simplicity of the proposed regression model makes it perfectly suitable as a baseline in addition to the convenience for applications such as the virtual try-on. To improve the repeatability of the results of our baseline and the competing methods, we provide guidelines toward standardized body measurement estimation&#8230;<br></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1srp178/warning_dont_get_gptbrained/?utm_source=share&amp;utm_medium=mweb3x&amp;utm_name=mweb3xcss&amp;utm_term=1&amp;utm_content=share_button">Warning: Don&#8217;t get GPT-brained [Reddit]</a></strong></p><p>At my last role we had to move fast, so we relied on an LLM to help with a lot of the thinking and coding for us so we could focus on the business use case and managing meetings and stakeholders. The role was heavy on project management as well as development, research, and deployment so basically doing everything While I got good at scoping projects and managing them, my technical skills totally deteriorated in less than 1 year. It&#8217;s scary going back to problems I know I can solve and but have some brain fog when getting to the answer. If I could have gone slower, had more time to thinking about modeling/coding than I probably wouldn&#8217;t feel like this Don&#8217;t get GPT brained. You&#8217;ll have to crawl out of that pit eventually. Like technical debt but for your brain&#8230;</p><p></p></li><li><p><strong><a href="https://arxiv.org/abs/2604.21691">There Will Be a Scientific Theory of Deep Learning</a><br></strong>In this paper, we make the case that a scientific theory of deep learning is emerging. By this we mean a theory which characterizes important properties and statistics of the training process, hidden representations, final weights, and performance of neural networks. We pull together major strands of ongoing research in deep learning theory and identify five growing bodies of work that point toward such a theory: (a) solvable idealized settings that provide intuition for learning dynamics in realistic systems; (b) tractable limits that reveal insights into fundamental learning phenomena; (c) simple mathematical laws that capture important macroscopic observables; (d) theories of hyperparameters that disentangle them from the rest of the training process, leaving simpler systems behind; and (e) universal behaviors shared across systems and settings which clarify which phenomena call for explanation&#8230;<br></p></li><li><p><strong><a href="https://arxiv.org/abs/2509.09892">What do the fundamental constants of physics tell us about life?</a><br></strong>In the 1970s, the renowned physicist Victor Weisskopf famously developed a research program to qualitatively explain properties of matter in terms of the fundamental constants of physics. But there was one type of matter prominently missing from Weisskopf&#8217;s analysis: life. Here, we develop Weisskopf-style arguments demonstrating how the fundamental constants of physics can be used to understand the properties of living systems. By combining biophysical arguments and dimensional analysis, we show that vital properties of chemical self-replicators, such as growth yield, minimum doubling time, and minimum power consumption in dormancy, can be quantitatively estimated using fundamental physical constants. The calculations highlight how the laws of physics constrain chemistry-based life on Earth, and if it exists, elsewhere in our universe&#8230;<br></p></li><li><p><strong><a href="https://github.com/entropich3atdeath/gpr-thermodynamic-hardware/blob/main/GPR_ThermodynamicHardware_v1.ipynb">Gaussian Process Regression (GPR) &amp; Physics-driven Computing</a></strong><br>Gaussian Process Regression (GPR) is a non-parametric, Bayesian approach to regression. Unlike parametric models (like linear regression) where we find a distribution over parameters, a Gaussian Process defines a prior <strong>distribution over functions</strong>&#8230;<br></p></li><li><p><strong><a href="https://opendatastructures.org/">Open Data Structures - An open content textbook</a></strong><br><em>Open Data Structures</em> covers the implementation and analysis of data structures for sequences (lists), queues, priority queues, unordered dictionaries, ordered dictionaries, and graphs&#8230;Data structures presented in the book include stacks, queues, deques, and lists implemented as arrays and linked-lists; space-efficient implementations of lists; skip lists; hash tables and hash codes; binary search trees including treaps, scapegoat trees, and red-black trees; integer searching structures including binary tries, x-fast tries, and y-fast tries; heaps, including implicit binary heaps and randomized meldable heaps; graphs, including adjacency matrix and ajacency list representations; and B-trees&#8230;<code><br></code></p></li><li><p><strong><a href="https://stormatics.tech/blogs/postgresql-is-not-slow-your-queries-are">PostgreSQL is Not Slow. Your Queries Are.</a><br></strong>A field guide to the seven things that are actually making our database feel slow and how to stop blaming the wrong suspect&#8230;Culprit #1: The Missing Index&#8230;Culprit #2: The N+1 Query, Death by a Thousand Cuts&#8230;Culprit #3: Stale Statistics, The Planner Is Flying Blind&#8230;Culprit #4: The Query That Runs Fine, Alone&#8230;Culprit #5: EXPLAIN ANALYZE Exists, Use It&#8230;Culprit #6: Connection Exhaustion, PostgreSQL is Full&#8230;Culprit #7: Reporting Queries Running on Production&#8230;<br></p></li><li><p><strong><a href="https://globalresearchspace.com/space#7.02/-4.771/61.204/-52.6/30">An interactive semantic map of the latest 10 million published papers</a></strong><br>I built a map to help navigate the complex scientific landscape through spatial exploration. How it works: Sourced the latest 10M papers from OpenAlex and generated embeddings using SPECTER 2 on titles and abstracts. Reduced dimensionality with UMAP, then applied Voronoi partitioning on density peaks to create distinct semantic neighborhoods. The floating topic labels are generated via custom labelling algorithms (definitely still a work in progress!). There is also support for both keyword and semantic queries, and there&#8217;s an analytics layer for ranking institutions, authors, and topics etc&#8230;<br></p></li><li><p><strong><a href="https://perthirtysix.com/how-the-heck-does-shazam-work">How The Heck Does Shazam Work?</a></strong></p><p>How audio fingerprinting and a connect-the-dots trick lets Shazam identify a song in seconds&#8230;By throwing away almost everything and keeping only a handful of landmark peaks, a noisy 5-second clip from a coffee shop becomes a set of coordinates precise enough to pinpoint one song out of millions. Recognition, it turns out, is mostly an exercise in ignoring the right things&#8230;<br></p></li><li><p><strong><a href="https://rtichoke.netlify.app/posts/generating-correlated-random-numbers.html">Generating Correlated Random Numbers in R Using Matrix Methods</a></strong><br>Generating random data with a specific correlation structure is a common need in statistical simulation. In this post, let&#8217;s walk through how to do it in R using matrix decomposition methods&#8230;<br></p></li><li><p><strong><a href="https://jcarroll.com.au/2026/04/17/schotter-plots-in-r/">Schotter Plots in R</a></strong></p><p>Translating things between languages reveals how each language approaches different design trade-offs, and I believe it&#8217;s a useful exercise. Having something to translate is the first step. I found a plot I wanted to generate (Georg Nees&#8217; &#8220;Schotter&#8221; computer-generated art from 1968 which shows a grid of squares which get increasingly displaced in position and rotation), and some code that reproduced it, so off we go!&#8230;<br></p></li><li><p><strong><a href="https://mfatihtuzen.github.io/posts/2026-04-16_timeseries_stationary/">Why Most Time Series Models Fail Before They Start</a><br></strong>A practical look at stationarity and transformations with real CPI data in R&#8230;Many time series models fail before they even begin. Not because the software crashes. Not because the code is wrong. But because the data entering the model violate one of the most important assumptions in time series analysis: stationarity&#8230;The goal is simple: show why raw time series levels often mislead us, what stationarity really means, and why transformations such as differencing and log-differencing are not cosmetic tricks but conceptual necessities&#8230;<br></p></li><li><p><strong><a href="https://kieranhealy.org/blog/archives/2026/04/25/hourly-subway-station-flows/">Hourly Subway Station Flows</a><br></strong>Pie charts are bad, as any fule kno. We&#8217;re not as good at judging relative differences between angles and areas as we are at judging relative differences in lengths on a common baseline. This is especially true when we have more than two things to compare at the same time. So, as a rule, you shouldn&#8217;t use them. You should figure out some other way of viewing your data instead. On the other hand, I just made 424 animated pie charts because if you&#8217;re going to break a rule you should break it good and hard&#8230;</p></li></ul><p>.</p><div><hr></div><h2>Last Week's Newsletter's 3 Most Clicked Links</h2><ul><li><p><strong><a href="https://science.nasa.gov/specials/your-name-in-landsat/">Your Name in Landsat &#128752;&#65039;</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/analytics/comments/1sqwb5l/ceo_cancels_bi_tooling_replaces_it_with_ai_breaks/">CEO cancels BI tooling, replaces it with AI, breaks everything [Reddit]</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/datascience/comments/1srxbjb/anyone_else_paranoid_using_ai_for_analysis/">Anyone else paranoid using AI for analysis? [Reddit]</a></strong></p></li></ul><p>.<br>* Based on unique clicks.<br>** Please take a look at last week's issue #648 <a href="https://datascienceweekly.substack.com/p/data-science-weekly-issue-648">here</a>.</p><div><hr></div><h2>Cutting Room Floor</h2><ul><li><p><strong><a href="https://perthirtysix.com/how-the-heck-does-gps-work">How The Heck Does GPS Work?</a></strong></p></li><li><p><strong><a href="https://www.dbos.dev/blog/benchmarking-workflow-execution-scalability-on-postgres">Does Postgres Scale?</a></strong></p></li><li><p><strong><a href="https://www.reddit.com/r/MachineLearning/comments/1sx3p40/how_do_you_test_ai_agents_in_production_the/">How do you test AI agents in production? The unpredictability is overwhelming. [Reddit]</a></strong></p></li><li><p><strong><a href="https://flovv.github.io/loyalty-programs-revisited/">Do Loyalty Programs Actually Create Loyalty?</a></strong></p></li><li><p><strong><a href="https://guillaumepressiat.github.io/blog/2026/04/logrittr-re">logrittr: A Verbose Pipe Operator for Logging dplyr Pipelines</a></strong></p></li><li><p><strong><a href="https://redwallanalytics.com/posts/2026-04-19-a-data-driven-survey-of-mlb-franchise-management/">A Data-Driven Survey of MLB Franchise Management - A search for factors leading to playoff success</a></strong></p></li></ul><p>.</p><div><hr></div><h2><strong>Whenever you're ready, 3 ways we can help:</strong><br></h2><ol><li><p><strong>Go deeper each week (paid subscription)</strong><br>Get 3 additional posts per week designed to help you:</p><ul><li><p>Statistics &#8594; understand the math behind ML</p></li><li><p>AI Agents &#8594; build with modern AI tools</p></li><li><p>Career &#8594; become more valuable at your job</p></li></ul><p><strong>&#128073; <a href="https://datascienceweekly.substack.com/subscribe">Upgrade for $10/month &#8212; cancel anytime</a><br></strong></p></li><li><p><strong>Looking to get a job?</strong><br>A practical guide to landing your first (or next) data science role, based on thousands of reader questions.<br><strong>&#128073; <a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">Check out our </a></strong><em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide">&#8220;Get A Data Science Job&#8221;</a></strong></em><strong><a href="https://www.datascienceweekly.org/data-science-guides/data-science-getting-started-guide"> Course</a></strong><br></p></li><li><p><strong>Promote your organization/project/event to ~68,500 subscribers<br></strong>Sponsor this newsletter and reach a highly engaged data science audience (30&#8211;35% open rate).<br><strong>&#128073; Reply to this email to learn more</strong></p></li></ol><div><hr></div><p>Thank you for joining us this week! :)</p><p>Stay Data Science-y!</p><p>All our best,<br>Hannah &amp; Sebastian</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://datascienceweekly.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Data Science Weekly Newsletter is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>