{"type":"rich","version":"1.0","provider_name":"Transistor","provider_url":"https://transistor.fm","author_name":"Regular Programming","title":"About Data Pipelines","html":"<iframe width=\"100%\" height=\"180\" frameborder=\"no\" scrolling=\"no\" seamless src=\"https://share.transistor.fm/e/59b05f05\"></iframe>","width":"100%","height":180,"duration":2616,"description":"Lars dove into data pipelines, and emerged bearing arrows and wishing for a lot fewer copies.\nWhat is there to think about regarding data pipelines, what is interesting about them?\nWhich tools are out there, and why might you want to use them?\nWhy all this talk about making fewer copies of data?\nWhat does Lars' current ideal pipeline look like, and where does Elixir fit in?\nLinks\nMatt Topol\nApache Arrow\nLarge language models\nVector search\nBigQuery\nsed\nAWK\njq\nReplacing Hadoop with bash - \"Command-line Tools can be 235x Faster than your Hadoop Cluster\"\nHadoop\nMapReduce\nUnix pipes\nDirected acyclic graph\ntee - to \"materialize inbetween states\"\nApache Beam\nApache Spark\nApache Flink\nApache Pulsar\nAirbyte - shoves data between systems using connectors\nCronjob\nFivetran - Airbyte competitor\nApache Airflow\nETL - Extract, transform, load\nDesigning data-intensive applications\nStream processing\nEphemerality\nData lake\nData warehouse\nThe people's front of Judea\nDBT - SQL-SQL batch-work-thingy\nSQL with Jinja templates\nSnowflake - data warehouse thing\nScala\nBroadway\nOban - \"robust job processing for Elixir\"\nDashbit\npandas - Python data library\nAPL\nArrow flight\nGRPC\nDataFusion - query execution engine\nPolars - \"DataFrames in Rust\"\nExplorer - built on top of Polars\nVoltron data\nThe Composable Codex\nPyarrow - Arrow bindings for Python\nQuotes\nI've been reading a lot about data pipelines\nWhat's so special about data pipelines?\nThere's a lot of special tooling\nThere's a lot of bad, bad tooling\nLess than optimal tooling\nConverging on something biggerlk\nHe got me eventually\nAll of your steps in one bucket\nWhat tools do you associate with data?\nI inherited a data pipeline\nBashReduce\nIterate on the L and the T\nThe modern data stack\nAnd then you demand more work\nNo unnecessary copies\nBarely a copy\nReconnecting with my Python roots","thumbnail_url":"https://img.transistorcdn.com/C_-AHQqsYbdliFKrXsdN7TWYO9_tZENNrKOb2TIA9xc/rs:fill:0:0:1/w:400/h:400/q:60/mb:500000/aHR0cHM6Ly9pbWct/dXBsb2FkLXByb2R1/Y3Rpb24udHJhbnNp/c3Rvci5mbS9zaG93/LzE5MjQ1LzE2MTg5/MzM2ODUtYXJ0d29y/ay5qcGc.webp","thumbnail_width":300,"thumbnail_height":300}