mysql: create table data (token_a varchar(10), token_b varchar(10));
Query OK, 0 rows affected (0.05 sec)
mysql: insert into data (token_a, token_b) values ('A', 'A');
Query OK, 1 row affected (0.01 sec)
mysql: insert into data (token_a, token_b) values ('A', 'B');
Query OK, 1 row affected (0.00 sec)
mysql: insert into data (token_a, token_b) values ('B', 'B');
Query OK, 1 row affected (0.00 sec)
mysql: insert into data (token_a, token_b) values ('B', 'A');
Query OK, 1 row affected (0.00 sec)
mysql: select * from data group by token_a;
+---------+---------+
| token_a | token_b |
+---------+---------+
| A | A |
| B | B |
+---------+---------+
2 rows in set (0.00 sec)
Note here the value we get for "token_b" is based on whether or not "A" or "B" were inserted first. The second "token_b" for each "token_a" (as well as any number of other rows that might follow it for that "token_a") is just discarded.The scary thing is that I semi-regularly come across applications in Very Important Industries that have large amounts of SQL that rely upon this behavior of "picking any old row" for you, rather than selecting a MAX() or MIN() of some column and then joining to a subquery of the GROUP BY + aggregate....because joining to a subquery in MySQL also performs like crap.