For scalar functions of a scalar variable, a less drastic option than nonstandard analysis (and perhaps one that’s more useful for bridging into more advanced stuff) can be as follows:
For any expression y involving x [*], denote by dy the linear part of y|x+t - y|x (that is, y evaluated at x+t less y evaluated at x). For y=f(x), that’s just a fancy way of saying f'(x)t, of course, but the intent is that where y was an “x-dependent scalar”, dy is an “x-dependent linear function” (without a constant term, as it is the convention in most settings outside of high school).
The baby version stops at that: by our rules, dx = t, df(x) = f'(x)t = f'(x)dx, f'(x) = df(x)/dx, the not-a-proof for the chain rule becomes an actual proof except everything is assumed differentiable (where a textbook one would only need differentiable inputs), etc. Of course, at this stage it seems slightly miraculous that nothing ends up t-dependent, but it is what it is. (If you want, you can imagine that the symbol t is “private” to the previous paragraph so it’s not allowed to escape to “the user”, but then you need to prove that it actually doesn’t.)
The adolescent version unholsters linear algebra:
While of course every one-dimensional (real) vector space (i.e. a line with a chosen zero point) is R in disguise, there are multiple choices for what the disguise is (differing by a multiplication by a constant), and you might not know which one to prefer. (This is what choosing a basis in a one-dimensional space amounts to.) Given such a space and two vectors u (whatever) and v (nonzero), denote by u/v the number such that u = (u/v)v (the coefficient of proportionality, aka the coordinate of v when e is the basis vector). Now you can prove, for example, that u/w = (u/v)(v/w), because it is so when you choose any (single-vector) basis and substitute for each vector its (only) coordinate.
(Can you make sense of uv/vw? Yes, as u⊗v / v⊗w, but tensor products are their own can of worms and probably overkill at this point.)
At each value of x, the space of linear functions (of the private variable t) is one-dimensional, so everything in the previous paragraph applies. When we write dy/dx and so on, we mean the things from there except we do them at each value of x (“pointwise”).
One of the things that this more advanced thinking gets you is that you can imagine how all of it generalizes to multiple variables. (Writing v/e_i for coordinates of v in the basis (e_i) is not common, but it does not not make sense—as long as you remember you can only “divide” by a basis, not by a single vector. Write out the coordinate transformation rules in this notation. The differential version will have you end up with df(x)/dx_i instead of the more common ∂f(x)/∂x_i, but again, that makes sense in context—note that, once again, the partial derivative wrt one coordinate on a plane depends on what the other coordinates on that plane are!)
The grown-up version just says I’ve been talking about the cotangent bundle in wishy-washy language. (An “x-dependent scalar”? What’s that? Does it taste good?) Hopefully it was still of some help.
[*] People do write df/dx where I would require df(x)/dx, and it’s convenient to do that, but I’m trying to avoid additional abuses of notation where possible.