It might be an interesting starting point for you: https://observablehq.com/@ajbouh/editor
174 karma · joined November 12, 2007
It might be an interesting starting point for you: https://observablehq.com/@ajbouh/editor
It will do things like separate out different kinds of test failures (by error message and stacktrace) and then measure their individual rates of incidence.
You can also ask it to reproduce a specific failure in a tight loop and once it succeeds it will drop you into a debugger session so you can explore what's going on.
There are demo videos in the project highlighting these techniques. Here's one: https://asciinema.org/a/dhdetw07drgyz78yr66bm57va
It's also the part of V2V that seems to be most powerful and least talked about.
I hope you're right and that once we establish best case utility, as a community we refocus on handling component failures more gracefully.
Although it's unclear if the second and third order system effects in our financial system will ever get that sort of treatment. So let's hope driving gets a bit closer to flying (and farther from wall street) in terms of attitude towards safety.
EDIT: clarity
Especially the associated threat modelling and engineering principles!
I'd rather just have to trust the safety standards of the manufacturer of my car (and to a lesser extent the cars I might directly collide with), not the safety standards of every vehicle within transmit distance.
That is, as a human I know that the drivers around me might have broken lights or misuse their signals.
I also assume that other drivers do not have my best interest at heart.
A reasonable V2V system should be expected to handle these scenarios just as easily as they handle wireless interference and packet loss.
V2V communication means BT exploits can become worms even more easily, spreading from phone to vehicle to vehicle.
My point is that V2V systems and research papers I've seen just don't make meaningful claims about safety. They instead make claims about convenience and efficiency, which are not substitutes for safety.
We know to test vehicles for crash safety before putting them on the road.
While I believe there are certainly ways to make use of V2V communication that increase overall safety, I haven't seen anything remotely resembling a crash test for V2V systems.
Creating a system that relies on the correct behavior of all components is not a recipe for safety or reliability. This is especially true as the number of components increases (this happens when a car exchanges messages with all the cars around it).
We need V2V systems that rely only on correct behavior of any component and allow for malicious behavior of some components.
The idea that a protocol version mismatch from some rolling deploy can cause injuries in cars made by other manufacturers is the sort of thing that I haven't seen a single person point out. How would something like this even be caught and debugged?
Advocating for an ecosystem where these things are likely but neither addressed nor considered is just plain irresponsible.
From https://knowledge.ni.com/KnowledgeArticleDetails?id=kA00Z000...
I have yet to see these sorts of considerations in any V2V communication system. The idea that my car might act on (or propagate) incorrect/sabotaged information that it receives from a "peer" is a terrifying form of fragility.
This sort of failure mode needs to be studied and addressed directly before any of these systems are deployed.
It should also be assumed that these new patterns of information propagation will create a huge financial incentive for people to sell after market modifications that exploit the trust models of these protocols.
It feels like your comment about the network and community being the most valuable part of program is probably right. Given that, is it possible to be a Pioneer, but not accept the money?
I'm very grateful for the work that this project has done and continues to do. Thank you!
The "industry specific deep learning" project is similar to something I'm working on right now. Though I'm not planning to charge for it.
To any folks here interested in this: Are you looking for a tool to get started in ML or for a resource to apply existing ML knowledge to a specific (possibly new) domain?
If any of our experiences or insights can help others in their own environments, all the better!
It was this sharing of artifacts that provided some of the impetus to use a sandbox, since a polluted output could poison the cache in hard to detect ways.
In terms of reporting on tests that run in parallel, we built a tool that specializes in exactly that. It collates output from parallel tests, it times out on tests that are hung, it makes sure the build system doesn't kill it if tests are too silent. It also tracks which tests have run against which versions of the codebase in the past and what their outcomes are. We use supporting tools to analyze test flakiness and understand when they are introduced. We have had a lot of success with this approach, as developers debugging weirdness across many tests is less miserable when they can use the same tools that CI does.
Critically, when bugs in those tools are discovered, developers can pinpoint and fix those bugs locally with reasonable ease. Deploying fixes to the test runner (or the logic that allocates workers for the test runner) is like any other change. No need to tinker with Jenkins (or buildbot, etc) config. No need to take the build system down to test that the change is correct. No need to bring up a test version of the build system and experiment with your change there.
We've gone to great lengths to make our system something that's a joy to work with and helped us be very productive across the many different environments we need to operate in.
It's tough to know how much detail is appropriate in comment threads like these. You're absolutely right that there's a lot that needs to come together to make something like what I've described work. I know because we pulled enough of it together to support our own large and heterogeneous projects.
It sounds like you have also thought about this problem a lot. Can you share more about the sorts of tests (language, test library, etc) you have? Perhaps we can break new ground where each of our respective experiences and intuition intersect.
CI runs the same exact build system (though with a few different options so the outputs are easier to during and after the build).
Passing CI is compulsory, as humans aren't allowed to release changes on our team. Humans may only do code review. If and when a change passes code review, it will be deployed automatically once it passes CI.
We use some of the same compute capacity that our CI system uses to scale test runners across many physical machines (though tests run against a pool of freshly cloned VMs using delta disks so we get a pretty big speedup and lots of control over the environment that tests run in).
There's a fascinating correlation between developer machines and build slaves. It's been my experience that needing to install system software of any kind on one usually leads to a headache later. We've gotten it down to just Xcode on OS X and almost just build-essential on Ubuntu.
So in spirit we do exactly what you're saying, we've just found a way to do it while using the same tooling on both CI and developer machines. We also demand that the build slave images are generated straight from install media and a fixed set of files (like those that install Xcode), so the only simple way to add dependencies (i.e. build tools or libraries) is via our build system. Use of apt, homebrew, etc is completely separate for our developers. And if they mess with the build system in a way that allows those files to leak in, the fact that build slaves are pristine means that their change will fail CI and never be deployed.
Does my explanation make sense? Happy to answer follow up questions. Also happy to be shown where our rigor is lacking :)
I have aspirations for QA (https://github.com/ajbouh/qa) to learn this trick, but it needs some lambda-specific smarts before it gets there.
When a build system can only be effectively invoked by CI/CD, it starts to pervert developer incentives. People need to check things in before they can be sure they work. They don't bother with tiny fixes because of the inertia. Flaky jobs get a quick rebuild, because reproducing a build failure locally is complex enough that they'd prefer to avoid it if they can.
Over time, these add up to a system that grows through accretion, which is the enemy of both agility and understandability.
Better is a build process that uses simple, reusable components that work equally well on developers' machines. These tools can be tested, refined, and replaced incrementally, using the same build processes that the rest of your code base does. You can do this without needing coupling your build processes to the specific way(s) that Jenkins (and company) model builds or their configuration.
There are an increasing number of build systems that encourage squashing these bugs. The resulting build outputs are simpler and are often more portable. They're also easier to reason about. That translates to simpler deployment, simpler operations, and fewer edge cases to debug.
IMHO, the most promising answer to the 5 minute limit is finer granularity and better caching of dependency inputs.
My point isn't that people should be asking for what you're building. My point is that you should think about how you'd know if people want it before you worry about making it.
That said, I regret adding that comment here, as it detracts from my larger point about encouraging YC to use more constructive ways to help their founders course correct.
Glad you mention the "make something people want" motto :). I've recently realized that it's actually wrong. Well, backwards.
A more instructive formulation is: "find something people want and make it." Better to acknowledge the value of the search (and confidence in the finding) before encouraging the making.
I have seen many founders struggle to understand when the time to grow is. Pressure to raise deepens the struggle. On the topic of growth vs retention, of the hundreds of founders I have met or read about, I have been most inspired by the actions of Ooshma Garg (of Gobble). The quality of her service is second to none. She's the only founder I've ever met that has completely shut down her product because it wasn't good enough, only to reboot it to better quality, retention, sales, and growth than ever before.
In my experience (flawed and narrow as it is), the best way to course correct is to issue a recall. Treat bad advice like bad goods. Post the criteria where those affected will recognize themselves and make needed reparations.
As a community, we need to better celebrate founders that are doing it right. Not mock those of us still figuring it out.
If YC wants to get the word out about prioritizing product market fit above growth, it should elevate Ooshma's perspective, along with other founders that have made similar gut-wrenching decisions to invest heavily in retention while sacrificing weekly growth targets.
Overall I'm glad you wrote and posted the essay. And I appreciate any humility you may have felt while drafting it.
This essay, while it discusses a really important idea, comes across as saying: "Foolish, fashion-focused founders! Clearly retention comes first! We would never imply that you should focus on growth before your product had adequate retention!"
But that's exactly what I've seen YC partners do. Recommend that companies focus on growth, because that's what YC's definition of a startup is: a company that grows very quickly. Having seen both sides of the curtain, the essay leaves me with a greasy, queasy feeling.
Am I missing something here?
Sam, are you changing your standing advice from "focus on growth" to something else? If that's what this post signals, please take ownership of the change and spell it out.
It's poor form to imply that misguided founders (and their devotion to a fad) are driving the growth zealot craze when you've had a hand on the wheel for years.
I first tried Gobble when I saw a trial card over at Zombie Runner on California Ave. Back then, Gobble delivered fully cooked dinners at 7pm on weekday nights. I used Gobble about 3 nights a week for a few months. Dinner went from something I had to figure out, to the highlight of my evening. We'd get to eat something new and delicious a few times a week without playing out local restaurants or the handful of dishes that we'd normally cook. Gobble actually improved my relationship with my girlfriend.
At the time, my primary complaint was that I had to reheat the food and eat out of a takeout container.
About a month ago, Gobble switched to delivering these meal kits. Having just met Ooshma in person, I (naturally) gave her a really hard time about this. Cooking takes too long and I have other things to do.
There were fewer choices than before, I needed to wash a pan (sometimes two, if pasta was involved), and I needed to actually get off my computer to help make dinner happen.
That was 15 meals ago.
I still think about the first Gobble dinner kit I made in back August. I actually woke up the next morning thinking about how good perfectly the fresh mozzarella balanced the spicy kick of the chili flakes.
Dinner-Kits Gobble is the best Gobble yet. Now I actually eat dinner like a real person. Sometimes dinner takes a little longer than 10 minutes to prepare, end to end, and sometimes it needs more than one pan, but it ends up tasting so good that I feel stupid complaining about either of those things.
Out of 15 meals, about 8 of them have been wish-there-were-more-leftovers-amazing about 5 of them were good enough order again, and a couple were just ok. All of them were worth way more than the $12 and ~10 minutes of effort.
I have never had a problem with Gobble that a message to Ooshma didn't fix.
Yesterday I got an email from Gobble with 3 invites for 100% free boxes. I've got two left. I think you still need to sign up to use it, but I'm pretty sure you can cancel immediately. If you want one of the invites, first come, first served.
[1] Unless a double espresso and a nice piece of bread counts as a meal.