31 karma · joined October 24, 2022
The problem is, CSVs don’t scale. No indexing means every search is a full scan. No structure means every query is brute force. A 5GB CSV isn’t just 5GB—it’s 15GB in RAM once it’s loaded, maybe more. If you don’t have the memory, your system starts swapping, and everything slows to a crawl. Sorting? Painful. Joins? Basically impossible. The tools we use weren’t built for this, but we keep using them anyway because, well, what else is there?
Analytics evolves across four levels: descriptive (what happened), diagnostic (why it happened), predictive (what will happen), and prescriptive (what to do about it). Each level requires different tools and capabilities, from data cleaning and interactivity to model explainability and operational integration. Delivery methods—real-time, batch, embedded, or ad-hoc—must align with the specific job, whether monitoring systems, diagnosing issues, or predicting trends. Success lies in fit: understanding users, the job at hand, and constraints like scale or compliance. The best analytics systems solve targeted problems exceptionally well, while the worst attempt to be one-size-fits-all.
This same transformation is possible for data apps. Right now, building even simple data apps is complex:
Setting up connectors to pull data Writing SQL to clean and model it Engineering a frontend and backend Deploying and maintaining the stack
The result? Data apps today are reserved for the most critical use cases. No one spends $50K to build a temporary dashboard for a weekend festival or a hyper-local app for a single restaurant. The ROI isn’t there.
But what happens when building data apps costs 1/100th of what it does today?
A restaurant manager could quickly analyze food waste patterns, broken down by dish, time, and chef. A festival organizer could track foot traffic, vendor sales, and bathroom wait times in real time. An individual sales rep could spin up a live dashboard to track their own deal pipeline.
What new behaviors and ecosystems would emerge if building data apps is this fast and cheap?
At first, we set out to solve a problem that seemed, at first glance, straightforward. Businesses have data locked in different systems: CRMs, product analytics platforms, billing tools, and more. To answer even basic business questions, they needed to extract and combine data from these silos.
Our approach was ambitious. We wanted to build a tool that handled every step of the data pipeline — data ingestion, modeling, schema mapping, metric definitions, and visualization — all through an intuitive and simple no-code SaaS interface. Our vision was to empower analysts, ops teams, and business users to do what once required a full data engineering team.
But, as we built and iterated and onboarded use-cases, we uncovered deep structural flaws in this approach. These learnings have been pivotal in shaping our current belief that low-code/no-code solutions are fundamentally misaligned with the challenges of real-world data complexity.
We learned that Code-first tools, with the right scaffolding, allow you to embrace complexity without being overwhelmed by it. With AI copilots like Preswald, code-first architectures, and unified control planes, we can finally create data stacks that are both powerful and approachable.
1⃣ Ingestion -- With comment-driven code, we connected to the USGS Earthquake Hazards Program API. Preswald generated the necessary Python code to ingest data with dlt and store it locally in a postgres instance.
2⃣ Transformation -- Preswald surfaced schema definitions and pre-suggested SQL transformations like seismic hotspots and earthquake frequency. Want to analyze trends by region or time? A few clicks, and you're there.
3⃣ Orchestration -- We configured a pipeline to refresh the data every 5 minutes and validate schema changes. Real-time data? Covered. Schema safety? Automatic.
4⃣ Visualization -- In seconds, Preswald generated the code for a Python-based interactive app—complete with a map to visualize seismic activity. Whether in VSCode or localhost, insights were live with minimal friction.
From API connection to live visualization, this was built entirely with code, in a fraction of the time it would typically take. Preswald automates the repetitive parts of building data apps while keeping developers in the driver’s seat.
Notebooks thrive because they solve one problem incredibly well: exploratory data analysis.
Low Barrier to Entry: Open a browser, type Python, and you're analyzing data. No setup, no fuss. Immediate Feedback: Write code, see results instantly—a dopamine loop for problem-solving. Ecosystem Ready: Seamlessly use tools like pandas and matplotlib. Storytelling in Context: Combine code, graphs, and narrative for sharable insights. But their success hardcoded limitations.
In today’s data workflows—ingestion, transformation, modeling, deployment, monitoring—notebooks falter:
Sequential Execution: Code runs top to bottom, making modularity and reuse painful. Collaboration Bottlenecks: Version control and multi-user editing are headaches. Deployment Gaps: Notebooks stop at analysis. Production requires messy rewrites. Notebooks weren't designed for modern complexity. Instead of forcing them to fit, the real question is: What comes next?
Why SWE-Bench Falls Short SWE-bench evaluates LLMs on real-world software engineering tasks by using GitHub issue–pull request pairs from popular repositories.
Success is measured by the quality, reliability, and scalability of data workflows, not just correctness of code.
Data engineers solve complex systems-level problems involving constant data motion, quality maintenance, and continuous evolution. Treating these disciplines as equivalent is doing data engineering a disservice.
DE deals with raw, messy, and large-scale data that must be cleaned, transformed, and made accessible.
i.e. handling data quality, schema evolution, and compliance, not just code correctness.
DE focuses on automating and orchestrating multi-step workflows, which requires understanding task dependencies, scheduling, and retries.
Edge-cases! DE must handle schema drift, missing values, malformed records, and outliers, which are rarely part of SWE workflows.
DE is less about writing application logic and more about making data usable, accessible, and reliable. Data engineers don’t just write code. They manage workflows that operate across multiple systems, handle changing requirements, and scale with data growth. A DE benchmark needs to reflect these realities.
One might argue that text-to-SQL benchmarks are a step toward evaluating LLMs for data engineering tasks. While useful, text-to-SQL falls far short of the needs of DE benchmarks for several reasons:
Text-to-SQL only handles querying structured data. Data engineering is so much more: it involves transforming, orchestrating, and making sense of chaotic, mixed-format data.
Lack of Pipeline Context. DE isn’t more than single queries. it’s about creating end-to-end workflows that deliver business value.
Handling real-world problems like schema drift or malformed records
Data engineers work with Spark, Airflow, dbt—not just databases. Generating a SQL query is child’s play compared to orchestrating a complex Spark job across petabytes of data.
Reliability, scalability, and optimization are make-or-break factors for DE. These are entirely missed by simple SQL benchmarks.
In short, text-to-SQL is just one piece of the DE puzzle. Evaluating DE copilots requires a broader, pipeline-focused benchmark.
A DE-bench would simulate real-world DE workflows, evaluating LLMs on their ability to solve practical, pipeline-oriented problems. Here’s how it might look:
Dataset Sources:
Use public datasets (e.g., NYC Taxi data, Kaggle datasets).
Generate synthetic data to simulate edge cases (e.g., missing values, schema drift).
Data Ingestion: Load raw data from APIs or files into a database or data lake.
Data Transformation: Normalize data formats, remove duplicates, and compute aggregates.
Pipeline Orchestration: Create an Airflow DAG for a multi-step ETL pipeline.
Schema Management: Migrate data between schemas safely.
A DE-bench would provide a structured, objective framework for assessing LLMs on real-world DE tasks, ensuring that these tools are reliable, efficient, and robust. It’s time we hold DE copilots to the same high standards as their SWE counterparts.