Union Types have been accepted for PHP 8
wiki.php.net
wiki.php.net
javascript has too many hidden ways to shoot yourself in the foot, and just hasn't evolved for developer happiness the way php and ruby have.
[1] https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Let me advice you to benchmark it "accuracy" on your machine setup, don't rely what is on that page.
Make sure the error with --jit-verbose=1 which will show whether it uses MJIT correctly.
ruby --jit-verbose=1
- N-Body single core
Ruby 2.6.3 6:22.00s
RubyJIT 2.6.3 3:58.18s
PHP 7.3.xx 4:10.90s
…
JIT success (511.7ms): initialize@nbody.rb:14 -> /tmp/_ruby_mjit_p5804u105.c
JIT success (397.5ms): block in offset_momentum@nbody.rb:68 -> /tmp/_ruby_mjit_p5804u106.c
JIT success (607.0ms): block in energy@nbody.rb:50 -> /tmp/_ruby_mjit_p5804u108.c
JIT compaction (53.6ms): Compacted 111 methods -> /tmp/_ruby_mjit_p5804u111.so
Successful MJIT finish
real 6m4.201s
user 6m42.813s
sys 0m4.041s
$ time /opt/src/php-7.3.11/bin/php -n nbody.php 50000000
-0.169075164
-0.169059907
real 5m24.915s
user 5m24.808s
sys 0m0.020ruby 2.7.0dev (2019-11-23T07:06:30Z master b563439274) [x86_64-darwin18]
gtime -v /usr/local/bin/ruby --jit -W0 nbody.rb 50000000 -0.169075164 -0.169059907
Command being timed: "/usr/local/bin/ruby --jit -W0 nbody.rb 50000000"
User time (seconds): 249.30 System time (seconds): 0.58 Percent of CPU this job got: 100% Elapsed (wall clock) time (h:mm:ss or m:ss): 4:08.97
---
PHP 7.3.11 (cli) (built: Oct 24 2019 11:29:52) ( NTS ) Copyright (c) 1997-2018 The PHP Group Zend Engine v3.3.11, Copyright (c) 1998-2018 Zend Technologies with Zend OPcache v7.3.11, Copyright (c) 1999-2018, by Zend Technologies
gtime -v php -n nbody.php 50000000
-0.169075164 -0.169059907
Command being timed: "php -n nbody.php 50000000" User time (seconds): 248.02 System time (seconds): 0.49 Percent of CPU this job got: 99% Elapsed (wall clock) time (h:mm:ss or m:ss): 4:09.82
but i'd pick php over js any day! (probably even over python/django, unless data mining is a core concern)
Yet there are still many ugly sides to PHP, and Laravel illustrates most of them. There is so much magic that IDE can't follow: some classes have a `__call()` magic function that redirect methods calls to other instances. Some functions return values of varying types, with no common interface.
I've worked with several PHP frameworks, and Laravel is by far the worst. Its awful documentation plays a big role in this (no real reference doc, just a tutorial ; no links to classes or function in the doc ; the API doc is a joke ; the acclaimed "laracasts" are useless for serious work). The fact that this framework is dominant in the PHP community is worrying.
I don't think it's about preferring one approach or the other.
A good library/framework should have both instructional documentation and reference documentation. They have different use-cases and are not interchangeable.
And then the trouble starts. Remember what properties your models had? No? Well so doesn't your IDE. There is just too much magic going on to keep things maintainable.
If you like a framework like Laravel you should go with Symfony instead. But don't use annotations. Keeps things separated so you and your colleagues can find routing and database information at logical places instead of all over the place in classes.
Not ideal - since it’s not enforced in anyway (the same as local scope type doc declarations). But can be useful.
We use Symfony at work and we turned off the annotations package, use XML mapping for our models.
Of course, Symfony and Doctrine do have some frustation points.
I use Symfony for quite some time and at one point I stopped using Doctrine data mapper. The DBAL and the query builder are enough.
I work for a web development company and We have been successfully using Yii for years now, including for medium size projects (in terms of LOC and exposure).
Just my two cents
Laravel is lauded by many for its fantastic documentation. It's not comprehensive, but it's still pretty extensive, easy to read, and the gaps are only for niche situations.
As far as the API docs, what is your complaint about them?
I disagree with you on this point. I worked with PHP for two years on some legacy systems, as well as some Laravel-based systems. While Laravel yields results that are worlds away from PHP build in the 90's, I still wasted dozens of hours because of poor constructs that the language is unfortunately coupled to. I documented each instance where I wasted a significant amount of time due to issues like scoping issues, closure problems, arrays-that-are-arrays-sometimes-but-maps-other-times, etc.
PHP has gotten better. But there are so many good choices that provide much better tooling, guarantees, etc. Examples: Ruby, Go, and Elixir. Heck, even Perl is more sane about its data types in some ways by enforcing consistent comparisons with `==` vs. `eq` and much more robust data structures.
I agree that every language has it's quirks. But I think PHP has so many that it seriously gets in the way often. I don't think it offers any significant velocity gains over using, say, Ruby.
a = []
a.foo = 'bar' // works!
b = {}
b[0] = 42 // also works!
I happen to think these designs are both INSANE, but even if you like them, PHP has no advantage here.[1] So how do arrays have numerical keys? Well... actually they don't. I was trying to make JS look better than it is in my original comment. Actually, array keys are stringified versions of numbers, and values are automatically cast to string when you put them in indexing braces.
Object.keys([1, 2])
--> ["0", "1"]
As you can imagine, this leads to gotchas when you assume that a JS object can have, say, a date as a key b = {}; b[new Date()] = 10;
Object.keys(b)
--> [
"Sat Nov 09 2019 23:15:58 GMT-0800 (Pacific Standard Time)"
]https://www.stefanjudis.com/today-i-learned/property-order-i...
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
Arrays are objects of class Array which have special literals, override the toString() method, and update the non-enumerable "length" property on certain operations.
Try calling Object.keys([1, 2, 3]), or even better try Object.getOwnPropertyNames([1, 2, 3])
In Zend/zend_types.h[1] of PHP source:
typedef struct _zend_array zend_array;
typedef struct _zend_array HashTable;
struct _zend_array {
...
};
That being said AFAICT, HashTable and zend_array are used interchangeably throughout the source. I am not a C programmer, but I did write a couple of PHP extensions and that was my general understanding. Perhaps it is a compatibility issue or just used to abstract types differently in various areas of the C API.Check out [2] for a deeper understanding of how arrays are handled internally in PHP.
[1]: https://github.com/php/php-src/blob/php-7.3.11/Zend/zend_typ...
[2]: https://nikic.github.io/2014/12/22/PHPs-new-hashtable-implem...
"An array in PHP is actually an ordered map."
god, the time i spent going paranoid debugging a stray "undefined array index 0" just to find out
array_filter(
is_uppercase,
['a', 'B', 'c', 'D']
)
returns [1 => 'B', 3 => 'D']
and not ['B', 'D']
like every other `filter(...)` i've used...EDIT to be fair, i see the rationale – having the index of the filtered element is useful sometimes, and requires some contortions with the usual impl of filter. it's quite neat, because you get both the index and the elements! but it's just... surprising as the default, and PHP's conflation of maps and lists obscures it – all the docs say is "array keys are preserved", which makes sense in retrospect, but doesn't really jump out for something with this much impact
For example: I don't think anyone would seriously defend the idea that brainfuck is a reasonable language to write production code in.
From the original post: [quote] Supporting union types in the language allows us to move more type information from phpdoc into function signatures, ... [/quote]
Dynamic is like only using untagged unions and primitives.
Paraphrasing:
1. Dynamic and Static lang code look different
2. Dynamic lang code can use variables of undeclared type
Neither statement adequately answers the parent comment.
― Linus Torvalds
Discussion: https://news.ycombinator.com/item?id=4560334
After you optimize the data structures and have established relationships, worry about the code.
// I know nothing. Not sure what types would work best...
public function method($data)
// I know more
public function method($data array)
// More
public function method($data array|false)
// Even more
public function method($data array|false) : int
// Ah. Okay. I understand the problem domain and have proper data structures to solve it
public function method($data array|false) : int|float
Notice the type system never got in our way. We organically grew into the type system as our understanding of the data evolved.That said, PHP 7 has squeezed most of the performance from the language and the low hanging fruit is gone.
However, PHP 8 looks to bring a standard JIT to PHP [1]. I'd guess that once that happens Hack / HHVM may no longer have any advantage.
[1] https://hub.packtpub.com/php-8-and-7-4-to-come-with-just-in-...
True that you shouldn't expect large existing PHP projects to run. But, the syntax is close enough that for net new projects, it's essentially PHP.
The built in webserver, Proxygen, is also usable in production and supports TLS, http/2, etc.
Imagine the fun!
<?php
function callit(float $numI) {
$numB = $numI || ($numI/2);
echo gettype($numB)."\n";
$numF = ($numI/2);
echo gettype($numF)."\n";
return some($numB);
}function someType(int|float $numX):int {
echo gettype($numX)."\n";
return $numX;
}$ret = callit(2);
echo gettype($ret)."\n";
<?php declare(strict_types=1);
Your code would flag an error in the IDE and throw a TypeError exception.
I don't see the need for abstraction either especially with Postgres which has a very good type casting system. I can insert a float into an integer column without any trouble and can return the same. Postgres even has a great syntactical sugar of :: for casting such that value::int or value::numeric just works.
Finally if this is a true requirement Postgres fully supports domains which are custom data types. They are not difficult to deal with and could provide for a syntax which would handle that. You might need a little work but I can have a domain which would include the int type and float type as a single data type. This still requires some plumbing to create a true union however it's not that far off. The downside however is again the DB has to store in binary so a domain like that would have two values stored for a single value which would be less than optimal solution.
The true value for PHP programmers here is that we can get closer to type validation in a duck typed system. This provides multiple type validation of scalars as parameter to a function. Less interesting to Java, Python, Ruby developers where everything is an object except in a few cases. Scalars are still widely used in PHP and this feature allows for not sending "banana" as a parameter to a function, that in the past would just cast this to 1. If you think banana == 1 then you are going to have logic bugs.
It's possible to emulate union types in SQL with triggers (i.e. check that the field A and B cannot be NULL at the same time, but it does not work if NULL is be a possible value.)
An aside: I wish some relation database adopted sum types: I didn't thought about the implications, but doing 'create table foo ( bar Maybe integer );' and then 'select Some bar ...' would be cool (and maybe a cleaner way to work with NULL.)
That's not an aside, that's my whole argument--RDBMSes should support sum types for all of the same reasons that we should use RDBMSes in the first place: developers describe a data model and work against that while the database storage and retreival. Postgres enums and NULL are just special cases of sum types.
For one, because they need a fixed binary representation for the type, to persist it on disk. In a programming language yoi do things in memory, so you don't have that issue...
Still, you could have had union types, or even coerce everything as string, as SQLite does, but it would be bad for performance as they'd need an alternate representation.
Memory and consequently programming languages also requires a fixed binary representation. How you represent data is orthogonal to its storage medium--you can write application memory to disk and read it back in, no problem (e.g., swap).
> Still, you could have had union types, or even coerce everything as string, as SQLite does, but it would be bad for performance as they'd need an alternate representation.
The problem doesn't go away by moving support out of the database and into the application; it only makes it worse insofar as the application is limited in its optimizations. At the end of the day, the real world has sum types, applications use them, and they are encoded into databases--they simply aren't encoded _well_ and the database isn't giving you any correctness guarantees as it does for product types (i.e., structs, records, etc).
Already covered that.
Programming languages can save their data as unions or structs, and take the minimal hit to switch on the type.
DB's persisting data on disk can do the same but will take a much bigger hit.
>How you represent data is orthogonal to its storage medium--you can write application memory to disk and read it back in, no problem (e.g., swap).
The costs are not orthogonal to the storage medium however.
>The problem doesn't go away by moving support out of the database and into the application
On the DBs side, it does go away. The DB only has to guarantee what it says it supports (only store one specific type in a column). So for the DB implementors, that's a great invariant for their implementation ease and performance.
If you mean primitive types specifically, this already is far from the case, and it's great. One "type" is great, but types can be composed. Allowing types beyond primitive types has already been a blessing for me. Getting columns of type ARRAY[some other type] or MAP[type, type] is incredibly convenient. I don't feel so strongly about the JSON types that are entering into every db, but they're certainly supported and widely used.
Where? I didn’t see it.
> Programming languages can save their data as unions or structs, and take the minimal hit to switch on the type. DB's persisting data on disk can do the same but will take a much bigger hit.
Yeah, of course. Disk is more expensive across the board. Same applies for storing ints, but databases don’t punt on that. And anyway, the sum types still exist in the schema, they are just implicit, as hoc spectacles built on the fly by the user. So you’re still dealing with performance issues, but they’re worse.
> On the DBs side, it does go away. The DB only has to guarantee what it says it supports (only store one specific type in a column). So for the DB implementors, that's a great invariant for their implementation ease and performance.
This applies to every feature for every tool. You don’t have to solve the problem if you just put it on your users.
Of course, tools have charters, and the relational axiom is that users shouldn’t have to manage their own storage and retrieval layer, but rather they should declare a data model and interface with it and the RDBMS would make search and retrieval fast and correct. Sum types are necessary in data modeling, so it fits clearly and neatly into the charter.
"For one, because they need a fixed binary representation for the type, to persist it on disk. In a programming language you do things in memory, so you don't have that issue..."
The crucial difference I point is "persist on disk" vs "do it in memory", not in the "binary representation". Both running programs and DBs have one, but one absolutely needs to be persisted on disk, whereas live program memory doesn't.
>This applies to every feature for every tool. You don’t have to solve the problem if you just put it on your users.
There's also the fact that it might not be a problem just an easy cop-out from the user.
In which case it's better to force your users into the more formal and rigid structure, and have them rethink their model, than turn the DB into an "anything goes anywhere" store.
Right, I agree, and my point was it doesn’t matter. Disk vs memory is a red herring. The same principles apply to both and the fact that disk is slower applies as much to product types as it does to sum types. In fact, sum types are represented as product types, but the system enforces invariants about the structure.
> There's also the fact that it might not be a problem just an easy cop-out from the user.
That’s a nope from me. Tools exist to solve problems. If a tool purports to solve a problem but only does it halfway, it warrants criticism or observation.
> In which case it's better to force your users into the more formal and rigid structure, and have them rethink their model, than turn the DB into an "anything goes anywhere" store.
RDBMSs are literally forcing their users into a less formal structure. You can’t rethink your model and make them go away (they are fundamental data modeling primitives), you can only find ways to hack product types to represent them, but you have to do all the work to make them fast and you probably just have to give up on verifiable safety altogether.
And how do you get from “sum types” to “anything goes data store”? Are you sure you understand the debate?
However, if I remove that declare from your latter example, only one output is shown.
Edit: Nevermind, I get it.
However, as with other dynamically typed and interpreted languages, all of this is happening at run-time, so one of the biggest benefits of strong typing (type checking at compile time) doesn't apply.
In my opinion, the PHP distribution shipping an official type checker which does nothing other than verify type correctness based on the information in the code would be much more useful. Kind of like `php -l`.
ISO C and C++ doesn't provide tooling, nor should any language. Tooling is better handled by third parties and implementers. Language should be focused on application and execution of the language. Unfortunately for PHP the language and implementation are tightly coupled as is Java and Swift.
A well constructed modern PHP project using composer has the tooling needed to statically validate. Personally I use PHPStorm but I also use IntelliJ for Java and Android. There are others out there as well.
Autoloading used correctly is no more than Java using a package line or namespace.
I find no problem with TypeScript and others. I however have a 10+ million line code base of PHP that predates TypeScript and others and needs to get a viable transition path. This gets things closer with real type errors at runtime. That is way better than my mainframe COBOL counterparts who have no path forward at the level of modern code.
As for the "language shouldn't provide tooling" argument: You picked C / C++ as a positive example for this. Those are standardized languages which evolve at a glacial pace. For most of their use cases, this is a good thing. But I'd say PHP's faster evolution over the last decade was the right thing for that language. Other modern language projects seem to follow a strategy of a single standard implementation with extensive tooling pretty successfully (e.g. Go, Rust, Swift, ..).
Edit: Ah I believe there is a misunerstanding between posters here. The usage of types is not enforced, but when you use them their correctness is enforced at runtime.