How PHP's foreach works
stackoverflow.com
stackoverflow.com
> Arrays in PHP are ordered hashtables (i.e. the hash buckets are part of a doubly linked list)
And that iteration is done using a "internal array pointer":
> This pointer is part of the HashTable structure and is basically just a pointer to the current hashtable Bucket. The internal array pointer is safe against modification, i.e. if the current Bucket is removed, then the internal array pointer will be updated to point to the next bucket.
Which together require some complex copying rules to allow for some simple things like iterating over the same array in nested loops.
I'm not very familiar with the implementation details of many other languages with these constructs, but in python a `for` loop (which operates similarly to the described `foreach` loop in php) simply operates over an iterator, which have a well defined implementation [1]. I don't know about the implementation any deeper than that, however.
I'm curious how other languages implementation of foreach type constructs stack up and how the choice of implementation for the standard list/array datatype affects the interface.
I was doing something fairly simple, trying to extract values passed as named argument to a function and turning them back into simple C types (char * and int):
// capturing hash keys as zvals
zval **salt_hex_val;
zval **key_hex_val;
zval **iterations_val;
if ( // getting values
zend_hash_find(hash, "salt", strlen("salt") + 1, (void**)&salt_hex_val) == FAILURE ||
zend_hash_find(hash, "key", strlen("key") + 1, (void**)&key_hex_val) == FAILURE ||
zend_hash_find(hash, "iterations", strlen("iterations") + 1, (void**)&iterations_val) == FAILURE ||
// checking types
Z_TYPE_PP(salt_hex_val) != IS_STRING ||
Z_TYPE_PP(key_hex_val) != IS_STRING ||
(Z_TYPE_PP(iterations_val) != IS_LONG && Z_TYPE_PP(iterations_val) != IS_DOUBLE)
) {
php_error_docref(NULL TSRMLS_CC, E_WARNING, "Could not extract and check types on required values in hash: salt, key, and iterations.");
RETURN_NULL();
}
char *salt_hex;
char *key_hex;
if (Z_STRLEN_PP(salt_hex_val) != salt_length * 2 ||
Z_STRLEN_PP(key_hex_val) != key_length * 2) {
php_error_docref(NULL TSRMLS_CC, E_WARNING, "Key or Salt length incorrect.");
RETURN_NULL();
}
salt_hex = Z_STRVAL_PP(salt_hex_val);
key_hex = Z_STRVAL_PP(key_hex_val);
int iterations = (Z_TYPE_PP(iterations_val) == IS_LONG ?
(int)Z_LVAL_PP(iterations_val) :
(int)Z_DVAL_PP(iterations_val));
The part that I still don't understand (but that I figured out by trial-and-error) was why `zend_hash_find` takes a `void••`[1] as argument, which should actually be a `zval•••` cast as `void••`. What's the purpose of the triple pointer here? zend_hash_find(hash, "salt", strlen("salt") + 1, (void**)&salt_hex_val)
[1]: Imagine the • there is a star / asterisk.1) The innermost pointer is needed because Zend hash tables actually store a "zval* " (pointer to zval), not a zval directly. The zval is allocated separately, then its pointer is stored into the table.
2) The second pointer is needed because Zend tables internally malloc storage for whatever they store (zval* ) in this case, then access that data as a pointer. The "zval* " pointer is memcpy'd into the malloc'd area. This data is accessed through a "zval* * " pointer. This allows users of zend_hash_xxx to not only access the "zval* " pointer, but also change it.
3) In C, one way for a function to return a value is by passing a pointer to a variable that will store the result. Since zend_hash_find returns the internal "zval* * " data, you need to pass in a "zval* * * " pointer to a "zval* * " pointer that is the actual return value you want. Through this "zval* * " pointer, you can read and also change the "zval* " data stored in that hash table cell.
So, to receive zval ••, you need to pass zval ••• to the hash function. That's why the type in general is void ••, because generic hash (hashes are used for all kinds of things, not only zvals) stores void •, so to receive it you pass void ••.
Just for fun, there are places in the code IIRC where quadruple pointer can be found (see zend_fcall_info_args_save for example). Pretty rare case though, don't remember any place with five-times pointer.
I've always found that an interesting choice.
At my last job, one of our interview questions for an experienced PHP programmer was "What makes PHP arrays different or unique, from a computer science point of view?" Not a single one ever got anywhere close to the right answer (that they're ordered hash tables, not arrays).
Pretty sure this applies to by-value and by-reference semantics too.
There was no clear explanation of references and their relationship with variables and how they differ from what are usually called 'pointers' that clearly distinguished PHP's view of variables from every other object oriented programming language that primarily used reference semantics for objects.
It was only in the context of a discussion of how zvals work and the is_ref and the refcount fields that I could understand exactly what semantics were being used by PHP.
It's a surprisingly useful data structure sometimes. I'm not sure if conflating three major kinds of data structures into one was a great decision, though. Honestly the OP's weirdness is much less likely to be important than the mixed-integer/string key weirdnesses you're more likely to run into.
Simple. You want to find out if they spend their time on HN. There hasn't been a single discussion on HN about PHP that doesn't mention that arrays in PHP are hash tables. Not one.
If the candidate is not smart enough to get up and leave the interview after being asked such a terrible question, you don't want to hire them. The company's actual business model is to sell the resumes of people who are smart enough to reject working there to companies looking for good coders.
Edit: confusing use of "their".
I'm still curious what do you mean by "ordered" because the PHP arrays don't order their keys. The common rant previously mentioned aka having to use asort() for getting a proper array:
php > $arr = array();
php > $arr[2] = 2;
php > $arr[1] = 1;
php > var_dump($arr);
array(2) {
[2]=>
int(2)
[1]=>
int(1)
}The order may be unusual or even non-obvious, but it is predictable.
By extension, a pitfall of interview questions: you can't always assume you actually know the correct answer or even that there is a correct answer. Which may cause interviewees to start secondguessing themselves and make them appear worse than they are. You can basically send them into an infinite loop by an illformed question. "Does he mean that they aren't actually arrays? Well, he can't mean that, because there are other languages whose arrays are actually hash tables. And they aren't called 'associative arrays' for nothing. So what does he mean?"
I'm not sure why this is interesting.
The answer to the question is pretty deep though.
$array = array(1,2,3,4,5); foreach ($array as &value){...}
foreach ($array as & $ref) {
// Do something with $ref
} unset ($ref);
i.e. put the `unset()` call on the same line as the closing brace, forever "welding" it to that block. $array = array(1,2,3,4,5);
foreach ($array as &$value) {
echo $value; // some non-mutating action
}
$value = 'woops';
print_r($array); // ([0] => 1 [1] => 2 [2] => 3 [3] => 4 [4] => woops) foreach($array as $key => $value)
{
if(someCondition($key))
$array[$key] = someTransformation($value);
}
or whatever - ie. just use the key to modify the original array directly. It's much easier to read and see what's happening (in my opinion anyway).