For RoR, see every method call, parameter and return value in production
callstacking.com
callstacking.com
(revenueTarget / 8760) * resolutionTimeTarget * numIncidentsTarget * resolutionTimeTarget + numEmployeesTarget * avgEmployeeTargeRate
This means revenue lost is correlated to the square of the lost time and the cost from employees is a static yearly cost.There are a couple things wrong with it. The third * should be switched with a + and the last term need to be multiplied by the number of incidents.
(revenueTarget / 8760) * resolutionTimeTarget * numIncidentsTarget + resolutionTimeTarget * numEmployeesTarget * avgEmployeeTargeRate * numIncidentsTarget
Which if anyone at Call Stacking is here, just means changing o = (n / 8760) * e * t * e + i * r;
to o = (n / 8760) * e * t + e * i * r * t;
or more succinctly o = t * e * (n / 8760 + i * r)
I'm assuming that's minified, so numIncidentsTarget * resolutionTimeTarget * (revenueTarget / 8760 + numEmployeesTarget * avgEmployeeRateTarget)
Edit: With the correct math, the example is wildly different. It should be $37,277.81, not $87,991.23.This would be your reference implementation.
https://github.com/callstacking/callstacking-rails
I'm sure we could work out an arrangement.
jim@callstacking.com
Something I've been interested in is the performance impact of using https://docs.ruby-lang.org/en/3.2/Coverage.html to find unused code by profiling production. Particularly using that to figure out any gems that are never called in production. Seems like it could be made fast.
prepend_around_action :callstacking_setup, if: -> { params[:debug] == '1' }
Once the request completes, the instrumented methods are removed to remove the performance overhead.
Also it is only necessary because of the lack of type enforcement which means no code can be relied and on and all code has to be constantly inspected for new bugs. Ugh.
Imagine a million-line codebase. There are half a dozen suspicious methods with complex sets of if/else if/else statements. And each of those statements make subsequent method calls.
Determining that code path is a nightmare. Types won't save you.
Types absolutely do prove logical correctness. In fact that’s all they do. However they can only prove the correctness of logic that’s type encoded. If your program is a primitive soup, there’s not much logic for them to prove.
That's not even mentioning the logic bugs you can spot more easily as well.
You have a new engineer. Point them to the Call Stacking dashboard.
"Here's a list of all of our endpoints for our application.
Click on a trace, and you can see the relevant methods for that endpoint and the context in which they are called."
This will get your new engineers up to speed, much quicker.
The trace URL will also be outut via the Rails log.
And the local usage section is hard to read white text on light blue background in this safari browser.
1) In a large-scale production scenario, you typically do not have the data, nor the interaction flow, to reproduce the bug locally. The idea is that you enable Call Stacking on the fly, when needed. Turn it off when not needed.
2) Having multiple runtime captures of the same endpoint across two different deployments or time periods allows you to quickly compare for logic or data changes (argument values and return values are visible).
3) Commenting on individual lines of execution allows for the team to have a specific discussion surrounding logic changes.
https://github.com/callstacking/callstacking-rails/blob/599d...
The goal is to quickly be able to see just the important, executed methods for a given request.
E.g. you may have a 2,000-line User model, but Call Stacking allows you to pinpoint, "Oh, only these three methods are actually being called during authentication. And here are the subsequent calls that those methods make. And here's where the logic change occurred."
You enable the instrumentation with a prepend_before_action,
e.g.
prepend_around_action :callstacking_setup, if: -> { params[:debug] == '1' }
When the request is completed, the instrumented methods are removed (thus removing the overhead).
You have to enable it judiciously. But for a problematic request, it will give the entire team a holistic view as to what is really happening for a given request. What methods are called, their calling parameters, and return values, all are given visibility.
You no longer have to reconstruct production scenarios piecemeal via the rails console.
Is it waiting for execution to return to the controller method and polling the stack trace from there?
When I was at ScoutAPM, we built a version of this that was stochastic instead of 100% predictable. We sampled the call stack every 10-50ms. Much lower overhead, and it caught the slower methods, which is quite helpful on its own, especially since slow behavior often isn't uniform, it happens on only a small handful of your biggest customers. But it certainly missed many fast executed methods.
Different approaches for sure, solve different issues.
On-premise is an option.
Not sure if Django could use it, but used an earlier version in Flask and Pyramid.
jim@callstacking.com
I hope you find somebody to help! Thank you!
Nearly all the issues this shows you quickly are issues that static typing would prevent compile time, or type-hints would show you in dev-time.
I've been doing fulltime Rails for 12+ years now, PHP before that, C before that. But always I developed side-gigs in Java, C# and other typed languages and now, finally fulltime over to Rust. They solve this.
Before production. You want this solved before production. Really.
Manually reconstructing logistical errors based on a combination of user input and system data, are the most time-consuming issues to diagnose.
When your codebase is 500,000+ lines of code, which code paths are relevant for a given endpoint? What methods were called and under what context? How do we begin to reconstruct this bug?
These are the scenarios for which Call Stacking gives instant visibility to.
Subjectively it feels like that catches more like 30-60% of bugs.
So this thing is still useful but Berkes is right that you need it a lot less if you use better static types.
> which code paths are relevant for a given endpoint?
This is exactly the question that static types can answer... statically. You don't need a runtime log to find out.
You do need a runtime log to see the actual values though. So it's not like a debugger is completely useless in Rust. But I definitely reach for it much less than in other languages.
This tool seems to be able to display the relevant code paths. That sounds super convenient and useful. Do statically typed languages have tools to do similar?
Really? Are you counting errors where a value turned out to be nil when it wasn’t expected to be? Because that’s a type error. It’s not a type error that many statically typed languages fix (Java is notorious for null pointer exceptions) but it’s a type error that can, in principle, be fixed with static typing.
It’s staggering really. These could be solved with better programming practices, they’re not errors usually experienced programmers make, but they exist nonetheless.
An exception/error system, with checked exceptions or handling of errors via result monad/errors like in Go, would solve the vast majority of other errors.
Type systems also encourage other things like contracts on API endpoints and typed messages, serving as a early warning system when writing those bugs.
Purely business logic bugs are actually not common. I guess we haven’t had the opportunity to create those “higher order” bugs while wading through the others.
I feel like 60% of the errors I encounter is high for type errors.
I also feel like with testing, fixing those errors occupies less than 5% of my debugging time.
Note before I'm downvoted for my last comment, there are exceptions but looking at the average candidate applying for a php, nodejs position compared to one for rails and one for ocaml.
I’m writing that to address the “just not in ruby” remark[1] from earlier.
Typing is great, I am a fan. But seeing the values that are running through those types is not solved at compile time.
Every time that serialization gets the wrong value passed in but continues anyway, is a type error. Every time a database-record misses a value (NULL) but the app continues to run over that, is a type error. And so on.
When I look at my most stable apps' Rollbar or Sentry, the top 20 errors are nearly all errors that a type system (which does not allow |null, like Java's - ugh, useless) would've caught compile or dev-time. The very few non-typing errors that are then left are race-conditions and business-logic-bugs.
The latter are really the only bugs that I'm "fine" with, they come with the domain. Race-conditions are quite often deep down also typing issues - they surface as similar `Undefined method on nil` errors because some data is nil due to the racing threads. Something a typing system would partly fix - as we can see in Rust.