262 karma · joined July 10, 2012
Previously: Co-founder @ Shipyard (www.shipyardapp.com) Head of Data Services @ PMG Digital Agency(www.pmg.com)
Get to know me - blakeburch.com Email me - small.tea5390[at]fastmail.com --- meet.hn/city/us-Austin
Socials:
- linkedin.com/in/blakeburch
- x.com/BlakeBurch_
- github.com/blakeburch
Interests:
AI/ML, Data Science, Gaming, Music, Privacy, Technology
---
wonderfuldev_hyi6bgph3zuhrqzoxp0sb5s7
This one-sided process is dehumanizing due to a lack of visibility and transparency of the process. Inconsistencies in the process develop when the hiring managers are responsible for too much and just let things slip. I'm of the opinion that increased automation could actually make the process feel more humanizing.
Here's my ideal state I've envisioned in a platform. I'm not in a position to explore this right now, but I'd love to chat with someone who is.
1. Employers should be forced to define their applicant funnel stages. Each stage, including the initial application, should be associated with an expected timeline and an automated message.
2. Rejection types and corresponding automated messages should all be set up front. Standard picks would be "Did Not Meet Requirements" (location, sponsorship, etc.), "Bad Fit (Resume Review)", "Bad Fit (Screening)", "Bad Fit (Post-Interview)", "Did Not Respond", and "Position Filled/Closed".
3. Applicants should have a dedicated portal where they can see the status of their application. It should show the entire funnel, their current status in that funnel, and the expected timeline based on their position in the funnel. It should even show details like who has viewed your application, when, and how many times. Additionally, all communication between the applicant and the employer should be shown here in a consolidated chat view.
4. When an employer moves the applicant from one stage to the next (kanban style) the applicant automatically receives a message with clearly defined steps to engage with the new stage. This should include scheduling links, project uploads, etc.
5. Stage timelines should be treated as SLAs. Employers can set up automated reminders to ensure they meet their SLA timeline for each candidate. If the employer doesn't meet their SLA, the situation should be handled based on automated rules. If the SLA isn't met due to lack of response/scheduling from the applicant, automated processes can be put in place to follow up or close out the application. If the SLA isn't met due to lack of employer interaction, the candidate gets the choice to let their application expire (incentivizing employers to be on time) OR the candidate can be automatically moved to the next stage (if feedback is present).
6. Using all of this information, applicants should be able to see an employers average response time, timeline adherence, etc.
7. When the position is closed, all applicants currently in the funnel should receive the appropriate rejection message.
8. Through (hopeful) economies of scale, applicants would eventually be able to track multiple applications through a single portal on their side.
If you want to get extra dicey, you could have it where applicants are also able to see team comments/feedback so employers are forced to be very structured about their feedback and there are fewer rejections without a shared understanding. I sympathize with both sides here.
--- I'm not a saint - I've created my fair share of bad candidate experiences. But it's never out of malice. It's always from bad process or tooling. Would love to hear thoughts on this idea.
Although I don't think most organizations are blaming lack of actionable insights on the data size. It's the lack of prioritizing data usage over data accessibility. We need to be teaching data people business levers and teaching business people data levers.
Data should be a byproduct of an actionable idea that you want to execute. It shouldn't exist until you have that experiment in mind.
We've gotten to a point where the first and last step get skipped. Business leaders see other companies doing interesting things with data, so the answer must be "gather all the data"! Internal teams end up focused on gathering the data without the context of how it might be used.
We need to train data teams to not focus on the data as the product. Instead, they should be responsible for driving business actions. Gathering and cleaning the data should just a byproduct of that activity.
Unfortunately, I've found that many data teams focus more on making the data clean and available. They never drive the conversation about what actions are being taken with the data. That leads to them being treated as cost centers. Wrote a similar post about my perspective on it - https://bytesdataaction.substack.com/p/transform-your-data-t...
I'd love to chat about the space more with you if you're interested! Email in bio.
Shipyard is a data orchestration platform that helps teams ship data anywhere in minutes without requiring code. We've built the platform to allow teams to mix and match their own code + our low-code templates to build powerful data workflows that are rooted in engineering best practices. We believe that with the right tooling, any data practitioner can be a 10x engineer.
Our next big era is focused on dynamic content creation, integration expansion, 1-click deployment, and building data workflows with natural language + AI.
If you're passionate about the data ecosystem and building the best technology for data engineers and analytics engineers, we'd love to have you join our fast-growing remote team. We have the financial backing of a larger company and are continuing to grow rapidly. Our back-end technology stack is built on AWS/Docker/Terraform/Postgres/Redis/Go/Sanity with the front-end built in React.js/Redux/Ant Design. Our 100+ low-code templates are developed in Python.
We're currently looking for the following roles:
- Integration Engineer (Python) | Develop low-code, open-source templates that ingest and send data between other SaaS tools
- Sales Development Lead | Own outbound outreach efforts and develop relationships with top brands and consultancies
- Full Stack Engineer (Golang) | Help accelerate our product vision through application development
- Growth Marketing Lead | Own SEO and content strategy and find unique ways to scale awareness
- Data Advocate | Build out use cases and content for data practitioners to learn from
If you're interested, apply directly through https://shipyard.breezy.hr/
I would love to know what the early journey was like when you took the plunge to design and publish games (and what parts of that you regret). I'm sure there was some imposter syndrome along the way.
We were just talking last week about how we should create a feature to describe transformations you want in Natural Language that get compiled to pandas/SQL. Input data is everything associated with the original file/dataframe.
Visual transformation tools are typically limited and non-reproducible. If you could switch it around to be code-compiled but description-driven, that would open up new possibilities.
I'd love to chat if you're open to it. Email in bio.
May need to scope if it's worth updating our open-source connectors.
But the poor vs rich game ended up with the poor person going up to $1000 and the rich person going down below $100.
I recognize it's just chance... but it's funny that the results directly conflicted the author's point.
It's worth checking out Shipyard (https://www.shipyardapp.com). I'm the co-founder and designed it to be the easiest way to deploy scripts/workflows in the cloud. You can schedule any code to run our platform (native support for Python, Node.js, Bash) and build reusable templates.
Our free dev plan allows for 10 hours of free runtime per month which is plenty for most use cases. If you want more or need webhooks/API, that starts at $50/month.
Feel free to contact if you want to learn more. Email is in my profile.
If you're looking for a way to quickly automate dbt Core, our team at Shipyard (low-code, cloud-hosted orchestration) recently built out some guides to start running dbt Core in the Cloud for all of the major databases. Plus, you gain the ability to connect to other tools in your data stack, run Python scripts alongside dbt, and still not worry about infra.
Text Guides: https://www.shipyardapp.com/docs/data-packages/dbt-core/dbt-... Video Guide: https://www.youtube.com/watch?v=wdbEiPHfKS4&list=PLsy6kuGU_w...
However, building a data product myself (Shipyard), we really try to encourage the idea of "connecting every data touchpoint together" so you can get an end-to-end view of how data is used and prevent downstream issues from ever occurring. This raised a few questions:
1. If the vendor owns the process of when the data gets delivered, how would a data team be able to have their pipelines react to the completion or failure of that specific vendor's delivery? Or does the ingestion process just become more of a black box?
While relying on a 3rd party ingestion platform or running ingestion scripts on your own orchestration platform isn't ideal, it at least centralizes the observability of ongoing ingestion processes into a single location.
2. From a business perspective, do you see a tool like Prequel encouraging businesses to restrict their data exports behind their own paywall rather than making the data accessible via external APIs?
--
Would love to connect and chat more if you're interested! Contact is in bio.
Tested this out in the span of a few hours and got a solution up and running to download the video from Youtube, spit out the transcription and upload the resulting transcription file externally. We're still missing a piece to upload directly to YouTube, but it's a start!
As a part of this experiment, we built out some templates that will allow anyone to play around with Whisper in our platform. If you're interested in seeing it, we built a video for doing the process with our templates [2], or directly with Python [3].
Hope someone finds this useful!
[1] https://www.shipyardapp.com [2] https://www.youtube.com/watch?v=XGr4v3aY1e8 [3] https://www.youtube.com/watch?v=xfJpGgyUkvM
> The tool data engineers need to be effective in this new world does not run scripts, it organizes systems. 100%. You'll still need to run independent scripts, but today's data challenges focus on "how do I connect the stages of data operations together". Teams need to figure out how to connect data ingestion -> data transformation -> data visualization -> alerting and reporting -> ML model deployment -> metadata + catalogs -> data augmentation -> API actions.
The larger goal of orchestration is to prevent downstream processes from running if the data being processed upstream fails. Each stage could be performed with a series of scripts, a SaaS tool, or a mix. Each team is responsible for their own stages, but they need to know how their work connects to the larger picture so when something goes wrong, there's ownership and clarity that drives a quick resolution. Unfortunately, this still doesn't exist in most organizations because the current tooling isn't solving the orchestration and visualization of connected systems super effectively. It's instead enabling one-off, disconnected data processes.
Disclaimer: I built Shipyard (www.shipyardapp.com) to address many of these concerns of simplifying the ability to connect data tools and quickly automate and action on data.
Shipyard is building modern data orchestration that's ridiculously easy for data teams to use. As a cloud-native solution with over 60+ low-code templates and the ability to run your own code with zero-config in the cloud, we help teams deploy data workflows in a matter of minutes. We believe that with the right tooling, any data practitioner can be a 10x engineer.
If you're passionate about the data ecosystem and building the best technology for data engineers and analytics engineers, we'd love to have you join our fast-growing remote team. We have the financial backing of a larger company and are aiming to achieve rapid growth over the next year. Our back-end technology stack is built on AWS/Docker/Terraform/Postgres/Redis/Go with the front-end built in React.js/Redux/Ant Design. We work directly with other tools in the modern data stack through the use of SQL/Python.
We're currently looking for the following roles:
- Solutions/Integration Engineer | Own the development and strategy of our low-code templates with other data partners
- Sales Development Lead | Own outbound outreach efforts and develop relationships with top brands
- Full Stack Engineer (Golang) | Help accelerate our product vision through application updates
- Growth Marketing Lead | Own content strategy and find unique ways to scale awareness
If you're interested, apply directly through https://shipyard.breezy.hr/
Feel free to message me separately if you want to dive into specifics.
Shipyard is building modern data orchestration that's ridiculously easy for data teams to use. As a cloud-native solution with over 60+ low-code templates and the ability to run your own code with zero-config in the cloud, we help teams deploy data workflows in a matter of minutes. We believe that with the right tooling, any data practitioner can be a 10x engineer.
If you're passionate about the data ecosystem and building the best technology for data engineers and analytics engineers, we'd love to have you join our fast-growing remote team. We have the financial backing of a larger company and are aiming to achieve rapid growth over the next year. Our back-end technology stack is built on AWS/Docker/Terraform/Postgres/Redis/Go with the front-end built in React.js/Redux/Ant Design. We work directly with other tools in the modern data stack through the use of SQL/Python.
We're currently looking for the following roles:
- Solutions/Integration Engineer | Own the development and strategy of our low-code templates with other data partners
- Sales Development Lead | Own outbound outreach efforts and develop relationships with top brands
- Full Stack Engineer (Golang) | Help accelerate our product vision through application updates
- Growth Marketing Lead | Own content strategy and find unique ways to scale awareness
If you're interested, apply directly through https://shipyard.breezy.hr/
What I've found while building out Shipyard, a hosted lightweight orchestration platform, is that teams want something that "just works". Servers that "just scale". Observability that doesn't require digging. Notifications and retries that work automatically. Workflows that don't mix business logic with platform logic. Code and workflows that sync with git. Deployment that only takes a few minutes.
For the straightforward use cases, where you need to run tasks A -> G daily, with a bit of branching logic, Airflow is overkill. Yes, Airflow has a lot of great complex functionality that can help you down the road. But Airflow keeps getting suggested to everyone even if it's not best suited to their use case, resulting in lots of lost time and engineering overhead.
While I have definitely have bias, there are a lot of other high quality alternatives out there to explore nowadays!
Also, Cristophe Blefari has an excellent data newsletter. https://www.blef.fr/
And Modern Data Stack has a newsletter, tool information, Q&A www.moderndatastack.xyz
One of my biggest annoyances in the orchestration space is that teams are mixing business logic with platform logic, while still touting "lack of vendor lock-in" because it's open source. At the point that you're importing Airflow specific operators into your script and changing the underlying code to make sure it works for the platform (XCom, task decorators, etc.), you are directly locking yourself in and making edits down the road even more difficult.
While some of the other players do a better job, their method of "code as workflow" still results in the same problems, where workflows get built as a "mega-script" instead of as modular components.
I'm a co-founder at Shipyard, a light-weight hosted orchestrator for data teams. One of our core principles is "Your code should run the same locally as it does on our platform". That means 0 changes to your code.
You can define the workflow in a drag and drop editor or with YAML. Each task is it's own independent script. At runtime, we automatically containerize each task and spin up ephemeral file storage for the workflow, letting you can run scripts one after the other, each in their own virtual environment, while still sharing generated files as if you were running them on your local machine. In practice, that means that individual tasks can be updated (in app or through GitHub sync) without having to touch the entire workflow.
I'm biased, but it seems crazy to me that so many engineers are willing to spend hours fighting the configuration of their orchestration platform rather than focusing on the solving the problems at hand with code.
Shipyard is building modern data orchestration that's ridiculously easy for data teams to use. As a cloud-native solution with over 60+ low-code templates and the ability to run your own code with zero-config in the cloud, we help teams deploy data workflows in a matter of minutes. We believe that with the right tooling, any data practitioner can be a 10x engineer.
If you're passionate about the data ecosystem and building the best technology for data engineers and analytics engineers, we'd love to have you join our fast-growing remote team. We have the financial backing of a larger company and are aiming to achieve rapid growth over the next year. Our back-end technology stack is built on AWS/Docker/Terraform/Postgres/Redis/Go with the front-end built in React.js/Redux/Ant Design. We work directly with other tools in the modern data stack through the use of SQL/Python.
We're currently looking for the following roles:
- Solutions/Integration Engineer | Own the development and strategy of our low-code templates with other data partners
- Sales Development Lead | Own outbound outreach efforts and develop relationships with top brands
- Full Stack Engineer (Golang) | Help accelerate our product vision through application updates
- Growth Marketing Lead | Own content strategy and find unique ways to scale awareness
If you're interested, apply directly through https://shipyard.breezy.hr/
Shipyard is building modern data orchestration that's ridiculously easy for data teams to use. As a cloud-native solution with over 60+ low-code templates and the ability to run your own code with zero-config in the cloud, we help teams deploy data workflows in a matter of minutes. We believe that with the right tooling, any data practitioner can be a 10x engineer.
If you're passionate about the data ecosystem and building the best technology for data engineers and analytics engineers, we'd love to have you join our fast-growing remote team. We have the financial backing of a larger company and are aiming to achieve rapid growth over the next year. Our back-end technology stack is built on AWS/Docker/Terraform/Postgres/Redis/Go with the front-end built in React.js/Redux/Ant Design. We work directly with other tools in the modern data stack through the use of SQL/Python.
We're currently looking for the following roles:
- Solutions/Integration Engineer | Own the development and strategy of our low-code templates with other data partners
- Sales Development Lead | Own outbound outreach efforts and develop relationships with top brands
- Full Stack Engineer (Golang) | Help accelerate our product vision through application updates
- Growth Marketing Lead | Own content strategy and find unique ways to scale awareness
If you're interested, apply directly through https://shipyard.breezy.hr/
Shipyard is building modern data orchestration that's ridiculously easy for data teams to use. As a cloud-native solution with over 60+ low-code templates and the ability to run your own code with zero-config in the cloud, we help teams deploy data workflows in a matter of minutes. We believe that with the right tooling, any data practitioner can be a 10x engineer.
If you're passionate about the data ecosystem and building the best technology for data engineers and analytics engineers, we'd love to have you join our fast-growing remote team. We have the financial backing of a larger company and are aiming to achieve rapid growth over the next year. Our back-end technology stack is built on AWS/Docker/Terraform/Postgres/Redis/Go with the front-end built in React.js/Redux/Ant Design. We work directly with other tools in the modern data stack through the use of SQL/Python.
We're currently looking for the following roles:
- Data Community Advocate | Educate and assist data practitioners by creating content and representing us externally - Solutions/Integration Engineer | Own the development and strategy of our low-code templates with other data partners - QA Engineer | Own automated and manual QA to ensure application and template changes are user-ready - Sales Development Lead | Own outbound outreach efforts and develop relationships with top brands - Full Stack Engineer (Golang) | Help accelerate our product vision through application updates
If you're interested, apply directly through https://shipyard.breezy.hr/
- Automatic DAG generation based on dbt-QL declared dependencies.
- The structure of where (db/schema) and how (table/view/temporary) things are built is defined in a YAML configuration, not the individual SQL statements.
- Testing/documentation baked in.
Sure, you can manage every select statement as its own task, but it becomes pretty infeasible once things scale.
dbt can still be administered alongside all other E and L-type tasks. It's just a Python CLI wrapped around SQL SELECT statements.
1) Specialized tools reduce the amount of engineering overhead. As a business, I primarily care about time to value. If I can use specialized SaaS to get my data centralized, clean, and synced across my tools in a week, why would I want to spend months building all of these processes from scratch?
Sure, I lose control, visibility, and more... but I was able to deliver value 3 months ahead of schedule.
2) Existing tools like Airflow are highly technical to get started with. You can't just focus on building out scripted solutions. You have to set up and manage the infrastructure. You have to sift through the tool's documentation to understand how to effectively build DAGs. You have to inject your business logic with platform logic to make sure your code will run on Airflow.
Because the demand for data professionals is high and the supply is low, the technology ends up trying to offset the need for those highly technical skills in your organization.
Data teams I talk to can't turn to any single location to see every touchpoint their data goes through. They're relying on each tool's independent scheduling system and hoping that everything runs at the right time without errors. If something breaks, bad data gets deployed and it becomes a mad scramble to verify which tool caused the error and which reports/dashboards/ML models/etc. were impacted downstream.
While these unbundled tools can get you 90% of the way to your desired end goal, you'll inevitably face a situation where your use case or SaaS tool is unsupported. In every situation like this I've ever faced, the team ultimately ends up writing and managing their own custom scripts to account for this situation. Now you have your unbundled tool + your custom script. Why not just manage all of the tools and your scripts from a singular source in the first place?
While unbundling is the reality, this new era of data technology will always still have a need for data orchestration tools that serve as a centralized view into your data workflows, whether that's Airflow or any of the new players in the space.
(Disclosure: I'm a co-founder of https://www.shipyardapp.com/, building better data orchestration for modern data teams)