Julia Language Co-Creators Win James H. Wilkinson Prize for Numerical Software
sinews.siam.org
sinews.siam.org
“A programming language cannot be derived from first principles alone. Language design is applied psychology—the computer program is the ultimate human-computer interface. Sometimes you have to try a design out and see how people interact with it and iterate based on that real-world feedback.”
I wish more people viewed PL like this.
For example, when I explained signed and unsigned to him, since we already covered the concepts of groups of bits having 2^n possible values, he quickly understood that signed has just as many values as unsigned, but the upper half just become negative inverse of the lower half.
C evolved from actual needs, and the fact that it's still widely used today shows how incredibly well designed it was for practical but relatively general problem solving. The C committee's job is mostly just a matter of saying "no" to every new feature request, because C has proven to be versatile enough to work in an extremely diverse array of situations.
We've been using Lua's C API and the Win32 API[1] to make cool things, such as a primitive Love2d clone, and just last night he said "Wait, you're telling me this is the old way of programming Windows? It just makes so much sense and seems so well designed!" C isn't perfect, but if you start from where C started, and trace it until now, you'll learn all the reasons behind its decisions, and they make sense in practical situations.
Kind of like the difference between metric and imperial measurement systems. Metric makes sense in an isolated mathematical world (i.e. science), bu imperial units make sense when you're working with real physical objects (e.g. construction).
I wish less people abbreviated "programming language" to PL.
For a shorter read with many of the same highlights, this is a paper about the design and implementation of Julia (PDF):
https://julialang.org/publications/julia-fresh-approach-BEKS...
Think of it this way: A single threaded C++ or Rust program can do more work per second than a hypothetical GIL-free CPython with 64 cores...
That's very context dependant. With annotations, typical calculations compile to C code roughly equivalent to native. The more of the actual runtime you use, the more you call into cpython which can't be sped up this way. Cython compiled code will be somewhere between C and cpython in terms of speed, but putting a single number to it will be always misleading.
> That's very context dependant. With annotations, typical calculations compile to C code roughly equivalent to native. The more of the actual runtime you use, the more you call into cpython which can't be sped up this way. Cython compiled code will be somewhere between C and cpython in terms of speed, but putting a single number to it will be always misleading.
OP said "CPython", not "Cython". CPython is pretty reliably 100X to 1000X slower than C.
Maybe, if your heavy lifting functions are very short lived, releasing and acquiring the GIL is holding you back...
Btw, I'm not defending the GIL. Python has enough visibility and significance to deserve a good implementation. It's just that most people don't realize the slowness of CPython itself is a bigger impediment to performance than the GIL. Python should be in the same speed league as JavaScript, but there are too many things hindering progress.
Numerical code is always written in numpy which can drop the GIL.
[1] https://discourse.julialang.org/t/not-only-for-technical-com...
I could totally see someone like my dad using it (he's an aerospace engineer, not software), but I have friends who work in compsci in academia trying to evangelize the language.
Not trying to start a war here, but I'm curious what more-seasoned Julia vet might have seen that someone with only ~12 hours with the language wouldn't.
edit: as another comment mentioned, metaprogramming is also great!
I've done a lot of work in both Python with Scipy/Numpy, and Julia. Python is painfully inexpressive in comparison. Not only this, but Julia has excellent type inference. Combined with the JIT, this makes it very fast. Inner numerical loops can be nearly as fast as C/Fortran.
Expanding on the macro system, this has allowed things like libraries that give easy support for GPU programming, fast automatic differentiation, seamless Python interop, etc.
I was unaware that Julia was homoiconic...I'm somewhat of a Lisp fanboy so I might need to give the language another chance.
Errors from missing or conflicting methods tend to not happen much in practice.
Damn. That's a pretty clever trade off between dynamic and static types.
It's not about putting any language or community down.
By your argument, it sounds like we cannot talk about any advantages a new language may have over existing ones. If we stuck to that, how could we possibly make any argument over why one may want to select a newer language?
This is not easily possible in Python, especially not with high performance.
Languages are toolsets and programmers are always looking for appropriate tools for a given task. It seems crazy to demand that people not be able to talk honestly about their experiences using a given tool. Python doesn't have feelings, it won't be sad if you don't like it.
I am experiencing something of a sense of schadenfreude over defense of Python against a newcomer. 20 years on, Perl is still in widespread use, still unbeatable for its growing use cases. I don't wish to speculate where Python will be 20 years from now. I do know that Fortran will still be around. That's about the only sure thing.
It depends on what you understand by homoiconic:
https://stackoverflow.com/questions/31733766/in-what-sense-a...
(math "x = sin(a) * b ^ 2")
Lack of something like that turned me away from Lisp (LISP at the time) a (cough) long time ago, since I didn't like the impedance mismatch between Lisp code and math for numerics.
At any rate, Julia looks like a fine general purpose language, I'm going to see if I can contribute somehow. I'm hopeful that GC pauses can be avoided for large classes of Julia programs, but we'll see...
(Let's not count things like computer algebra systems written in Lisp that have some sort of actual math notation for input and output, just alternative syntax for programming in Lisp itself.)
As early as 1973, there was https://en.wikipedia.org/wiki/CGOL
That can actually be made to work today: http://abcl-dev.blogspot.com/2010/04/cgol-on-abcl.html
Instead, people prefer to say that julia code is just another (tree-like) data-structure in the language and it can be manipulated at runtime with functions or compile time with macros or at parse time with string macros and now with Cassette.jl[1] we can even manipulate the form of code that has already been written and shipped by other packages all with first class metaprogamming tools. It seems to me that even if Julia is not 'truly homoiconic', that we seem to get the touted benefits of homoiconicity to the point that it seems like an unimportant distinction.
(I also vaguely recall Julia describing its type system as "dependent" in way that goes against convention. Maybe they just liked controversy in the early days!)
Is it a better language, from a pure CS perspective, than Python/Scheme/pickyourfavorite? No, not really. If your job is writing web pages, Julia will be fine but it won't excel.
But if your job is writing trajectory planning routines for robots or climate simulations or so on, you aren't coming from those languages. You are coming from Matlab. Matlab has great libraries -- every numerical method under the sun. It has a very shallow learning curve. It's good for prototyping. But the actual Matlab language, as opposed to the toolboxes, is absolutely painful for writing programs longer than a few pages. As a numerical methods guy with a CS background, it literally makes me want to tear my hair out. It feels like writing in BASIC, circa 1989.
Julia can also have a shallow learning curve. If you want to prototype a couple of pages of numerical methods, it works nearly as well as Matlab does. But if you then need to expand that code into production, it can do that too.
Could you be thinking more of a "pure developer perspective" than "pure CS perspective" perhaps? Because if anything I find Julia's type system a lot more interesting from a computer science point of view than Python or Scheme.
If I want to get things done, I'd stick to Python for now though due to sheer momentum.
Yes, sure. Julia is an interesting language, CS-wise. My only point, inasmuch as I had one, was that there are other languages out there that have those same features.
"If I want to get things done, I'd stick to Python for now though due to sheer momentum."
I tried python, because you're right, it has the momentum and is the obvious first choice for numerical computing after you've had your fill of Matlab. Unfortunately, for a lot of the code I write, speed matters. In large part I'm prototyping my own numerical methods, not totally relying on numpy/scipy. The last time I tried python for numerical code development I ended up with a routine that was two orders of magnitude slower than it needed to be. I rewrote the entire thing in Julia, with little to no effort towards optimizing it, and it was immediately 50x faster than python and within a factor of two of what I needed. I put a couple of days of effort into optimization and I got it down to 2-3x faster than required.
You could argue that I should have gotten better at optimizing python, and maybe you'd be right. I'm no python expert. But in my limited interaction with the python community, the answer to faster python seems to be, at the end of the day, to write the parts of the code you really need to be fast in C and call them with the python FFI. Which is what numpy does, if I understand correctly. Whereas I can write that same code directly in Julia and it's generally fast enough.
For a lot of use cases speed doesn't really matter. But for robotics (my field) or self driving cars or machine vision or a lot of other embedded applications, it does matter.
Multiple dispatch is such a great way not to have to write boilerplate code for every method, somehow java seems very awkward in comparison.
Julia feels like a thoughtfully designed and carefully constructed language. I don’t havd the training to understand why exactly.
Matlab is (was?) the default tool for computational science, and Julia borrows lots from it. I'd say it's safe to put Julia in the same category with Matlab and Mathematica
But I can't imagine anyone would choose Matlab or Mathematica to build a backend...
>But I can't imagine anyone would choose Matlab or Mathematica to build a backend...
Julia is a clean, concise language that has so far mainly been focused on numerics. People are looking at using it in many other areas, for instance real-time robotic control: http://www.juliarobotics.org/
This means that at least some Julia folk are looking at ways of avoiding garbage collection during critical code sections...
Data center efficiency is a concern, and there is real interest in better languages for web applications. Rust and Swift both have projects exploring this area. Julia would fit well in that space, and in my opinion its syntax is much cleaner than either.
Are there any other resources like this to learn fundamentals of Math with Julia? ([Linear] Algebra, Calculus, Geometry, Basic Physics, Discrete Math).
Jokes asides, congratulations to Jeff Bezanson, Stefan Karpinski, and Viral B. Shah, well deserved!
You mean you'll say only zero thing?
Introducing TwoBasedIndexing.jl (https://github.com/simonster/TwoBasedIndexing.jl)
The point is, n-based indexing isn't really that big a deal.
Here's context : https://jontysinai.github.io/jekyll/update/2019/01/18/unders...
There's nothing else like it.
Also here's the julia equiv to dplyr: https://github.com/queryverse/Query.jl
the more general version of the elevator pitch is that being in an environment where all the code is written in one high-level language that is still fast is qualitatively different. you can combine things in all kinds of strange ways with surprisingly little work. I.e. the DiffEqFlux package takes about a hundred lines of code to glue together a GPU neural network library and some very advanced differential equation solvers.
realistically, it sounds like you are hitting all of R's strongest points currently, though. I suspect right now you would be pretty dissatisfied with Julia.
s = 0 for k in range(0, N): s += a[k]
Ok, I have more cores/CPUs, could I do it in parallel? Sure
v1 = Sum(0, N/2) v2 = Sum(N/2, N) s = v1 + v2
I even could do it on asymmetric cores (like most phone CPUs today)
v1 = Sum(0, K) v2 = Sum(K, N) s = v1 + v2
I could do a bit more complicated things, like calculating means. For that, I have to make a monoid
def mean(from, to, a): m = 0 for k in range(from, to): m += a[k] N = to - from return (m/N, N)
Here I'm returning tuple, and for that to be a monoid I have to state composition law:
def compose(M1, M2): A1, N1 = M1 A2, N2 = M2 N = N1+N2 return ((A1N1 + A2N2)/N, N)
If my identity is (0,0) tuple, it is quite easy to verify that indeed I have a monoid. Nice, simple, composable ranges.
For closed [1...N] ranges this is NOT nice, NOT simple, just plain fugly exercise. Sorry, looks like code samples are screwed up a bit, don't know how to fix it
@threads for i in eachindex(object)
and let the details about how many cores be written once, correctly, elsewhere?Sure, if you have single infinitely fast CPU. In real life, we have to decompose problem, run it in parallel and compose it back.
> Or @threads for i in eachindex(object) & let the details about how many cores be written once, correctly, elsewhere?
Here I'm explicitly talking about details, how it should be done, importance of composition, monoids etc. I understand, that if someone did it for you, you might not care - well, more power to you then. But some basics rules about Pyton ranges are pretty good:
1. [0...K) + [K...N) = [0...N)
2. Given range [K...N), number of elements in the range is N-K, in case of [0...N) number of elements is N-0=N.
Simple, elegant, composable. I can't say the same about Julia
[1...N] = [1...N+1) = [1...K+1) + [K+1..N+1)
So the real loss is your point (2) - which, I agree, makes the implementation much less elegant and simple.
you could make them work, but improper abstraction leaks out in so many fugly ways. F.e. in Python you could compose functions, not only ranges. What do you pass in? Simple, [0, len(a)). What you get out? Simple, len(a). So you could operate on ranges composing function calls like in FP. Even works for unknown beforehand sequences/streams, just count along the sequence how many events you processed, and return it as length, and it could go into composition function. With Julia you have to think what is passed and what is returned. Shall I always return len(a)? Or maybe len(a)+1? What to return in case of streaming events? Sometimes len(a) and sometimes len(a)+1 with tons of comments and warnings? It is not a good way to deal with all that and not a good way to build API. Could be done and probably was already done, sure. BUt simplicity, elegance and composability is missing in the base design.
function mean(a, l, r)
s = 0
N = r-l+1
for k in l:r
s += a[k]
end
s/N, N
end
function mean2(a)
N = length(a)
lr1 = (1,div(N,2))
lr2 = (div(N,2)+1,N)
v1, N1 = mean(a, lr1...)
v2, N2 = mean(a, lr2...)
N = N1+N2
(v1 * N1 + v2 * N2) / N, N
end
function compose(M1, M2)
v1, N1 = M1
v2, N2 = M2
N = N1+N2
(v1 * N1 + v2 * N2) / N
end
function mean3(a)
N = length(a)
lr1 = (1,div(N,2))
lr2 = (div(N,2)+1,N)
compose(mean(a, lr1...), mean(a, lr2...))
end
You have to think about what is passed in and out anyway. Arguments of elegance etc. are most of the time personal preferences and very biased.No, they are not. I'm talking about ability to build (and build upon) underlying algebraic structure.
ok, could you make a monoid out of your/Julia closed ranges? What would be a unit in this monoid? How would you make a null/empty range?
Say I have an index `idx` into a very large array and I want to select a symmetric window of `N` indices before and after that index. In Julia:
A[idx .+ (-N:N)]
In Python: A[idx-N:idx+N+1]
Python cannot even index lists by ranges (or other lists). And if I go out of bounds, well, it'll happily just give me back a list of a size I didn't expect.Heck, in Julia I can even easily determine what the OOB behavior should be. It's an error above, but I can easily change it. This, for example, will ensure no indices go out of bounds by repeating the final endpoint as necessary.
A[clamp.(idx .+ (-N:N), 1, end)]
Sometimes you want to work in terms of the fenceposts, and sometimes by the length of the fence. There certainly are cases where 0-based offsets are nice — or even arbitrary offsets. In those cases we have OffsetArrays.True for python lists, but not true for numpy arrays.
Could you make proper monoid out of Julia ranges? What would be a unit in that monoid? Could you tell me how empty range looks in Julia (in Python/C++ it looks trivial)?
> Python cannot even index lists by ranges (or other lists)
sure, Python has deficiencies as well.
C++ STL iterators/ranges are made in the same way, and for a good reason - composability first. It described very well in Alex Stepanov book (https://www.amazon.com/Mathematics-Generic-Programming-Alexa...), as well as how it helps when we move to parallel algorithms (http://stepanovpapers.com/p5-austern.pdf).
It is even more explicit in upcoming Ranges library in C++20
split(r) = r[1:end÷2], r[end÷2+1:end]
That'll happily recurse and eventually spit out empty ranges given an input like `split(32:56)`.A zealous adherence to one strategy or another will simply blind you to the ways in which the other might be easier at different times.
Thankfully the changes in modern hardware architectures have triggered many runtime devs to start having another look at their compiler backends, improving vector instructions support and access to GPGPU features, instead of sticking with just good enough for CRUD apps.
But when it's between waiting 1 day or 1 month for your results you start caring quite a lot
We're not going to have a computer that can solve the Schrodinger Equation for a brain anytime soon.
If anything it's the inverse:
Not only Moore's law has stopped working for several years now ([1]), but we have also greatly increased what we do with computers (e.g. the whole dataset of an 80s company could fit on a 1GB disk -- today we routinely need to process terabytes of data both streaming and offline -- see also 4K video, webasm vs simple websites of 1999, VR, AR, and so on).
But even if computational power increased continuously and became cheaper and cheaper it would only help if our data and processing needs remained static or increased at a significantly slower pace.
[1] https://steveblank.com/2018/09/12/the-end-of-more-the-death-...
https://www.extremetech.com/computing/256558-nvidias-ceo-dec...
https://interestingengineering.com/no-more-transistors-the-e...
Re: the "webasm vs simple websites of 1999" factor, I think that's a case of designers (and the programmers enabling them) self-inducing such problems. Most business sites don't actually need all those newfangled features and tools; plain old HTML5+CSS3 (plus maybe a little bit of JS here and there) is more than good enough.
Also in static languages, the set of available instance is closed and extending it requires approaches like extension methods now available in a couple of OOP languages.
Whereas multi-methods are dynamically dispatched taking into consideration the whole set of parameters, and can be defined after the fact. Meaning that they aren't necessarily written in the same module that defines the respective data structures, which opens the door to more extensible designs.
Taking the Alan Perlis' quote "It is better to have 100 functions operate on one data structure than 10 functions on 10 data structures." to the next level.
Because not only can you define all those functions that operate on one data structure, they actually apply to the tuple made by the dynamic type of all parameters at call site.
Naturally it also might make it harder to follow what actually happens with a given method dispatch.
This is my insight from having gone through "The Art of the Metaobject Protocol" and dabbling with multi-methods in Clojure, so I might not be 100% correct here.
Building on top of your remark, Julia's approach is more "call the visible implementation of this function that better matches the types of all given parameters".
And these are only two possible ways of doing OOP, there are a few other ways of approaching the ideas of writing extensible modular code with polyphormism.
In other words: object orientation != strong typing (even if they do tend to go hand-in-hand).
CLOS, Beta, SELF, Oberon (the first version), Component Pascal all have explored different ways of doing OOP.
Or that you have to say foo(x) instead of x.foo()?
The latter makes writing code using autocomplete less doable, but reads more naturally in many cases.
I can't think of compelling pragmatic reason to mind the former, but if you have one I'd like to hear it.
But the problem is not that big. Maybe the methods are right next to the struct, or if I search the codebase for the struct I can find which methods use it. Still, would be cool if Juno had some feature like "show which methods use this."
Just because Visual Studio supports Visual Basic doesn't mean I can't use it when I'm only interested in writing things in C#.
---
500 - Internal server error.
SQL Exception
Error Details
File
Error Index #: 0
Source: .Net SqlClient Data Provider
Class: 17
Number: 1105
Procedure: AddEventLog
Message: System.Data.SqlClient.SqlException (0x80131904): Could not allocate space for object 'dbo.EventLog'.'PK_EventLogMaster' in database 'DNN-PROD' because the 'PRIMARY' filegroup is full. Create disk space by deleting unneeded files, dropping objects in the filegroup, adding additional files to the filegroup, or setting autogrowth on for existing files in the filegroup. at System.Data.SqlClient.SqlConnection.OnError(SqlException exception, Boolean breakConnection, Action`1 wrapCloseInAction) at System.Data.SqlClient.SqlInternalConnection.OnError(SqlException exception, Boolean breakConnection, Action`1 wrapCloseInAction) at System.Data.SqlClient.TdsParser.ThrowExceptionAndWarning(TdsParserStateObject stateObj, Boolean callerHasConnectionLock, Boolean asyncClose) at System.Data.SqlClient.TdsParser.TryRun(RunBehavior runBehavior, SqlCommand cmdHandler, SqlDataReader dataStream, BulkCopySimpleResultSet bulkCopyHandler, TdsParserStateObject stateObj, Boolean& dataReady) at System.Data.SqlClient.SqlCommand.FinishExecuteReader(SqlDataReader ds, RunBehavior runBehavior, String resetOptionsString, Boolean isInternal, Boolean forDescribeParameterEncryption, Boolean shouldCacheForAlwaysEncrypted) at System.Data.SqlClient.SqlCommand.RunExecuteReaderTds(CommandBehavior cmdBehavior, RunBehavior runBehavior, Boolean returnStream, Boolean async, Int32 timeout, Task& task, Boolean asyncWrite, Boolean inRetry, SqlDataReader ds, Boolean describeParameterEncryptionRequest) at System.Data.SqlClient.SqlCommand.RunExecuteReader(CommandBehavior cmdBehavior, RunBehavior runBehavior, Boolean returnStream, String method, TaskCompletionSource`1 completion, Int32 timeout, Task& task, Boolean& usedCache, Boolean asyncWrite, Boolean inRetry) at System.Data.SqlClient.SqlCommand.InternalExecuteNonQuery(TaskCompletionSource`1 completion, String methodName, Boolean sendToPipe, Int32 timeout, Boolean& usedCache, Boolean asyncWrite, Boolean inRetry) at System.Data.SqlClient.SqlCommand.ExecuteNonQuery() at PetaPoco.Database.Execute(String sql, Object[] args) at DotNetNuke.Data.PetaPoco.PetaPocoHelper.ExecuteNonQuery(String connectionString, CommandType type, Int32 timeout, String sql, Object[] args) at DotNetNuke.Data.SqlDataProvider.ExecuteNonQuery(String procedureName, Object[] commandParameters) at DotNetNuke.Data.DataProvider.AddLog(String logGUID, String logTypeKey, Int32 logUserID, String logUserName, Int32 logPortalID, String logPortalName, DateTime logCreateDate, String logServerName, String logProperties, Int32 logConfigID) at DotNetNuke.Services.Log.EventLog.DBLoggingProvider.WriteLog(LogQueueItem logQueueItem) ClientConnectionId:84098483-739c-4cd8-bf62-38ff4b256361 Error Number:1105,State:2,Class:17
PS: And in regards to your last statement, maybe "end"s aren't seen very much anymore is because language designers' agree with me, not that somehow my opinion is influenced by the syntax of modern languages.
If you just think about it, that is much more likely than what you are proposing.
As to the benefits of not using `{}` for blocks, pairs of ASCII brackets are one of the most precious commodities in programming language syntax design. Wasting them on blocks is not ideal. For example, C++ was out of brackets to use for template parameters, and had to use `<>` instead—despite the syntactic clash with their meaning as less than and greater than operators. This has caused no end of parsing problems and irregularities in C++ and related languages and has only been made to work because they're all static languages with separate "expression context" and "type context" and inequality operators don't make sense in type contexts. In a dynamic language, there's only one context and you can't distinguish the use of `<>` as brackets from its use as inequality operators based on that. Julia uses `()` for function calls (pretty important), `[]` for array indexing (also pretty crucial), and `{}` for type parameters.
Using indentation is a fairly clever choice but not without its drawbacks. It causes lots of problems with cutting and pasting code. It means that your IDE cannot generally autoformat or autoindent your code for you. There are also many people who find the "trailing off into space" visual appearance of Python code unsettling and prefer the symmetry and closure provided by `end` or other block delimiters.
In regards to you the use of brackets, array indexing and function calls can use the same syntax: lisp and scala seem to manage just fine. There's a lot of unnecessary syntax in a lot of languages that doesn't really serve any purpose (besides historical).
I'm not trying to argue about any particular syntax. Just that, syntax should be as concise as possible and not include extra identifiers that confuse the eye.
As an aside, I also think that programming languages should probably use less words, and more symbols, as it lets people who are familiar with other spoken languages to more easily adopt them.
Also, using more keywords (like end), restricts the name space of variable names.
Yes, Matlab does this too. However, not syntactically distinguishing array indexing from function calls has problems in a language that provides as much array functionality as Julia does. The classic example in Matlab is that when the interpreter sees `a(b(c(end))))` it needs to dynamically look up whether `c` is an array or not to decide if the `end` refers to the length of `c` or not, if not then it has to check if `b` is an array or not, etc. That would be even worse in Julia since the question of "is it an array or not" can be somewhat fuzzy. It's perfectly possible to implement an array-like type that uses indexing syntax and does not subtype AbstractArray. In Julia, when you see `log(v[end÷2])` it's syntactically unambiguous that this means `log(v[length(v)÷2])` which is only possible because of the distinct syntax for array indexing and calling functions. There are other where this distinction is essential as well, such as broadcasting [1].
[1] https://docs.julialang.org/en/v1/manual/arrays/index.html#Br...
> I actually agree with you on the last part, I do find python's tab indention unsettling. I just think it's better than using "end"s.
Ok, well we made the opposite judgement—we found `end` preferable to indentation. When you create a language you get to pick syntax that appeals to you.
> I'm not trying to argue about any particular syntax. Just that, syntax should be as concise as possible and not include extra identifiers that confuse the eye. > As an aside, I also think that programming languages should probably use less words, and more symbols, as it lets people who are familiar with other spoken languages to more easily adopt them. > Also, using more keywords (like end), restricts the name space of variable names.
You may appreciate K and other APL derivatives, although they are not known for their readability. But they certainly don't favor English speakers over anyone else.
The best example is probably the differential equation solution type in Julia. When you have a solution `sol`, the interface gives you `sol[i]` as the values in the solution's time series and `sol(t)` the continuous solution of the differential equation. This are natural ways to describe both the discrete array-like and continuous function-like nature of an ODE's numerical solution, and that syntax wouldn't be possible without distinguishing a function call from indexing.
I suppose it could indeed be jarring if you're not already used to languages which use 'end' to terminate blocks, but after awhile your brain will just start to treat 'end' synonymously with '}' and all will be right with the world.
Maybe someone can justify why a 3 character word (end) doesn't really matter. But can someone give an argument on how it adds value versus one character (a bracket), or zero characters (visible at least), a tab?
The crux of the issue is the signal to noise ratio. The more noise there is (syntax), the more difficult it is to pick up the signal (the semantics).
begin some code some more code end
It's easier to read.
You can certainly argue that '}' is language-agnostic and therefore a better choice, and I wouldn't disagree with you. But 'end' is by no means an arbitrary choice for a character sequence that ends a block.
Yes, easy to write. Seriously, I can type short words like end just as quickly as I can type shift-}, two keys off in the extremes of the keyboard layout.
I think it gets a bad reputation from languages like bash that have a horribly inconsistent mix of `fi`, `end`, `esac`, and require opening words like `then` and `do`, which people can argue about where to place.
end worked great in pascal and matlab, and it works great in julia as well.
By sharing block delineation with variable namespaces, the language creates a whole bunch of problems (a block identifier is essentially a newline followed by a “end” vs the much simpler “}”).
Also, the more english centric a language becomes, the harder it is to learn for non-english language speakers.
Moreover, words are used for variable names, while symbols are disallowed is many/most languages. By having block delineaters share the variable namespace, you detract from the readability of the language (as well as restricting variable names as I mentioned earlier).
One might ask: Does any of this matter in the scheme of things? The fair and honest answer is: not really, it’s a minor syntactic difference. But I firmly believe using words to end blocks is still inferior, even if it’s not a big difference.
Block delineation in Julia doesn't require the newline. You can just throw a ` end` at the end of the line. Often you'll see folks put the `;` line delimiter there, too, to make things super-clear (as that's generally how you delineate multiple expressions on the same line), but it's not required for an `end`.
And while it is indeed a little sad that you can't use `end` as a variable name, we make the most of it by using it to also represent the last index in an indexing expression like `A[end]`.
So, yeah, not a big deal at all, but very few of your detractions even apply in the first place. :)
APL, AWK, COBOL, Fortran, Lua, Mathematica, MATLAB, R, Smalltalk, Wolfram
Because it is Mathmatics and have 1 be based on 0 makes no sense, you then have to switch between the two and it is easy to make a mistake. This is why I don't use Python and Pandas. I got burnt once and that was enough and switched to R. Sadly we are stuck with 0 based array in programming and due to a historical issue.
Although in Wirth inspired languages, actually the indexes are flexible, but by convention 1 gets used.
like now with everything 0 based having something like:
arr[offset1 * 32 + offset2]
would be the following if all my offsets would be 1-based and arrays would be 1-based:
arr[(offset1 -1) * 32 + offset2]
which is pretty arbitrary...
My issue is higher level languages Python, Java, C# there is plenty of things that just are complicated and doing arrays off 0 is one of them. Doing data science or statistics just makes it obvious that you have two different sets on numbers. 1 doesn't mean the same thing in every instance in your programming and the functional programming side of me hates that.
Personally I'm happy to switch between both, and they both have positives and negatives when I actually write code.
0 indexing is not incorrect, or incompatible with maths, it is just a different way of conceiving of lists/arrays by considering the index to more like the distance from the beginning. The first item begins 0 spaces from the start position.
That is why in that context array length is usually the highest index + 1, because length is the distance from the start position to the end of the list, whereas the index is the start position of the item in the list.
I will concede that 0 indexing is a little too close to the storage paradigm for arrays (i.e. a sequential list of items of n bytes) and that when so often these arrays point to more complex data-structures / objects, the distance doesn't make as much sense as the item number... but meh - I'd take Python generators / iterators at the expense of 0 indexing any day.
The problem is that one is indeed an indexing (numbering the contents 1...X...N and asking for item X, customers[X]), whereas the other is not, but is used as an indexing (e.g. customers[5] is not getting the 5th item but the sixth).
If I had to give a name to the concept you're talking about, it would be an "ordinal index".
Which is both the naive understanding of a numeric index in everyday life in general, and the most common case in mathematics (and most math/scientific software).
I really do with the 1 indexing had stayed around, since I feel like it's more natural to say the 1st element, instead of the 0th.
It's not the 0th element. It is the 1st element and has an offset of 0 from the beginning of the array.
Not saying it's better or worse, but that if you change the words you use to describe it, you may find it easier to use.
To my surprise, I just learned that ranges in Rust don't include the end number which felt incredibly stupid from a conceptual point of view (python too). But is syntactically useful when doing:
for x in 0..len(my_array) {
My preference would be that a range include the end number and the language provide more python-like iterators and list comprehensions. I am aware that there is a library that provides more python-like syntax, but that doesn't change the oddness of ranges in either language. I guess I just need to remember that the interval is open on one end. Which brings to mind the notion of using [0..n] or [0..n) in a language...
It should be indexed with its position then a[1], as we do in all other aspects of life (and in math), and not with it's offset.
You mean like elevators? You may be surprised to learn that elevators in America are 1-indexed (ground floor == floor 1), but in Europe they are 0-indexed (ground floor == floor 0).
In most European countries the ground floor is not considered the same thing as the others. As Wikipedia puts it: "In most of Europe, the "first floor" is the first level above ground level".
We consider and count as floors the "layers" above the ground level. In french for example, the ground level is called "rez-de-chaussée" and the floors above it "XX étage". The ground level is not an "étage" (in American terms, the "ground level" is not considered a floor, and is not counted among the floors).
For this reason, apart from elevator labels nobody calls the ground level "the 0th floor" when speaking/writing.
And even in elevators, O is just used as a convenient way to convey the ground level (since the buttons don't fit much) -- and not everywhere. In many elevators it gets it own designation (e.g. G or GRD or some other national variety).
It's "the ground level" (with different names per country) + "N floors".
A multi-floor building with ground floor + 3 floors can even named "3rd-storey building" in some countries (and the ground level is just taken as granted) -- e.g. une maison à deux étages in the UK --> 3 story building in the US.
I think you just refuted your own previous assertion that 1-indexing is universal in all other aspects of life.
So, no 0-based indexing here.
We call the house a "3-storey house" not because we zero index, but because the cardinality of it's floors (as europeans define a floor) is 3. In "3-storey building" for Europeans, 3 is the "length", not the last index.
It's no more zero-based indexing than naming relations as “sibling”,“first cousin”,“second cousin” is, just because some hypothetical culture that also uses 1-based indexing but counts siblings as first cousins might rationalize our system as starting with “zeroth cousins” where they have “first cousins”.
What they have merely names what Americans name as zero.
But it's not a stand-in for zero, for us, it's a different entity (a different thing).
A floor (etage, piano, orofos, etc) is a "layer of building where people live above the ground level layer".
So there's no zero-based indexing of floors, because the first item in the set of what are considered as floors is called "first" (e.g. primo in italian).
So, the thing one must understand before they say "same difference", is the other cultural difference: that most European countries don't consider the ground level "floor" to be a floor/story.
I.e, we don't consider the upper floors and the "ground floor" to be in the same set. Which is also why we don't count them when we say a "2 story house" (and we mean a house with 3 such layers, ground story + 2 stories).
Is there a zero-based indexing of the whole "heterogenous" set of ground level + floors?
Well, it's common to have the ground level designated with G or national specific names: BG, BV, E, IS, PB, PT, RC, S, P, PK, etc.
It's just that in some elevators 0 is a substitute for this -- which is probably due to the few common brands and parts being international and nobody bothering to change them (e.g. Otis). That's more a red herring than how Europeans think about floors or index them.
So the floors are numbered according to their distance above the ground level (whatever that's called). This matches my suggestion of treating zero-based indexing as a distance from the first element.
>> Is there a zero-based indexing of the whole "heterogenous" set of ground level + floors?
Yes there is! It's how the elevator buttons are labeled with the possible variation of using a label instead of "0".
The European elevator numbering scheme according to your own comments matches more closely with a zero-based indexing scheme than a 1-based. Or more precisely it matches my proposal to consider the numbers used to reference array elements as distances rather than indexes.
It is true that some elevators label the ground level 0, but this is usually when there are also negative levels (cellars).
The German word Stockwerk is even clearer. It describes the wooden planks which separates two levels. So if you live on the first Stockwerk you live one level above ground.
If I say `for x in range(10)` it iterates exactly 10 times, starting at zero (which is convenient for 0-index lists in python). It would be extremely counter-intuitive if it did [0,10] as it would iterate 11 times. I agree though that it would be much less ambiguous if you could just do `for x in [0..10)`, for example.
Current implementation requires the least amount of "off-by-one" corrections for most tasks, IMO.
Right and I expect that. The issue I have is that the values do not include 10. It turns out great in Python and Rust where things are 0-indexed, but that does not change the fact that the range construct itself is strange in not including 10.
I'm sure the notion of using mathematical notation where [1..10], (0..10], [1..11), and (0..11) all mean the same thing would drive the language parsing people nuts. Mostly because some of those involve unmatched [) or (]. But that's why we have a computer right?
https://doc.rust-lang.org/std/ops/struct.RangeToInclusive.ht...
This is why I think the terminology should be "indexed arrays" and "offset arrays" (and variables perhaps even called index/idx/i or offset/ofst/o as appropriate) instead of using the terms "0-indexed" or "1-indexed."
But I have a hard enough time getting out of the habit myself, and some people are going to hate using o as a variable, so....
Also it's usually a bad sign if you need to worry too much about what index you used, you might be better of using index sets with whatever structure you need (and only the structure you need). In program languages that support it generous application of ranges and generators tends to get rid of most complications.
1 + 1 = 2 always means two. When you are getting the 1 column we are getting the second column. This while makes sense when doing certain lower level programming this makes no sense with functional or higher level programming.
There is an interesting note in the Wikipedia page for "Dartmouth BASIC" which says the second edition "also allowed arrays to begin as subscript 0 instead of 1. This was especially useful for representing polynomials."
Regex and Binary indexes are zero-based, List and Tuple are one-based
iex> [:foo, :bar, :baz] |> Enum.at(1)
:bar
Not that it really matters anyway. Like with Erlang, in the vast majority of cases where I'd normally reach for querying a list element by index, pattern matching is the better / more idiomatic option. Compare: foo = list |> Enum.at(0)
bar = list |> Enum.at(1)
baz = list |> Enum.at(2)
rem = list |> Enum.slice(3..-1)
with: [foo, bar, baz | rem] = list
Both give you the exact same values of foo/bar/baz/rem, but the latter is arguably more readable and concise, and entirely avoids the 0 v. 1 debate.https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...
One-based indexing works just as well as 0-based. I've done a lot in the past in Fortran 90 and (regrettably) Matlab. Some things are very slightly harder, some things are very slightly easier. At any rate, this isn't Fortran 77. You shouldn't be fiddling around with indices much.
The bigger difference IMO is Julia uses Fortran-style column-major ordering, instead of row-major like C and Python. I actually
integer :: array1(5)
integer :: array2(0:4)
integer :: array3(5:10)
integer :: array4(-4:4)I'd love to see what happens to the tabs vs spaces debate once someone makes tabs that aren't an integer multiple of spaces wide.
So, that world was and is a long time ago. We just don't typically give a way for people to set tabstops in most environments, therefore it was chosen that tabs would advance a certain number of spaces.
I've actually found behaviour of "zero-avoidance" a good battle-tested heuristics to coding. A lot of numeric spaces can be mapped as to be >0 or >0<. Obviously (?) not all of them can, but it's helpful in reducing cyclomatic complexity in logic code, as well as to consumers so they can uniformly handle all falsey values.
Arithmetic is much safer without 0.
if you have array [a,b,c,d] array[0] is a which in computer is logical. why? because it stored in memory 'abcd' (forgetting endianness for sake of argument...) and the offset from the base pointer to a is 0. so it makes perfect sense for arrays to start at 0, as there is no offset from the base... there is not 1 item offset from the base :s if you find it confusing i would say learn how computers actually operate instead of arguing from some theoretical frame of mind which has nothing to do with how computers work. a computer isn't mathematics or some arithmetic machine.
just because arithmetic is safer without 0 doesn't mean it's logical for a computer to suddenly have different array indexing. array indexing on a computer has nothing to do with maths. just memory layout and pointers...
#addsfueltofire
I don't recall the details, as this was a long time ago, but there were definitely times I was glad C++ used 0-based arrays, and there were definitely times when 1-based would have made for cleaner code.
So I overloaded the () operator to make a(n), where a is an array, equivalent to a[n-1].
It looked a little odd at first if I mixed a[] and a() form in the same program, but it wasn't hard to get used to.
I doubt I would recommend this in general, but it worked well for this particular application.
- Naming a PL "Julia" is sexist and creepy
- The language doesn't add any substantial value to PL theory or programming methodology compared to say, Python
- The future of computer-human interaction is not in programming. Python will be good enough until we reach that point where we can use neural networks/AI to do most of the business tasks.
How? (I kind of assume that a big reason computer scientists would pull out that name without a specific referent in mind is the influence of the Julia set, named for Gaston Julia.)
> The language doesn't add any substantial value to PL theory or programming methodology compared to say, Python
Maybe; not all practical benefits in programming come from advances in methodology or PL theory.
> The future of computer-human interaction is not in programming.
I've been hearing that since the early 80s. When someone makes that future reality, we can discuss what it requires, until then programming is what we have and I'm glad it wasn't neglected for the last four decades, and I hope it won't be for the next however many it takes, either.
- I don't think Python added anything to PL theory or methodology either. At least Julia introduced multiple dispatch to a new audience.
- How is the name sexist or creepy?