https://www.afr.com/companies/financial-services/anz-the-fir...
716 karma · joined August 19, 2010
https://www.afr.com/companies/financial-services/anz-the-fir...
It was absolute hell. The key problem here is the waiting and uncertainty. You have the NIPT at 10w, but you can’t have the amniocentesis until several weeks later. When that came back fine, there were questions about whether it was a “mosaic” meaning only a small proportion of cells are effected. We were only really in the clear after the 20 week ultrasound.
That’s a lot of weeks to be consumed by wondering about whether to terminate the pregnancy, or wait it out for more information. I have a masters in bioinformatics (in genomics!) and my knowledge of stats and the science was next to useless in the face of these decisions.
I know of couples who simply couldn’t deal with this uncertainty and chose to terminate on the basis of this test alone.
Fortunately for us our child was fine and is a perfectly healthy 18 month old now, but I wouldn’t do the rare trisomy test again.
The "data mesh" is essentially the collection of these independent "data-products".
We already see management problems with self-service analytics like PowerBI, Tableau & Looker. Its too easy for people to create dashboards / reports that are subtly wrong and which cause confusion. There is a balance between empowering to build data products and centralised control. Too much empowerment of people who don't understand the right way to do something leads to a horrible mess of contradictory data. Not enough, and people can't effectively do their job. Governance and process is the key to finding the balance and enforcing it.
The issue with the data-mesh is that there isn't really any great tooling to support the management or development of data products, or a data-mesh generally. I am sure this will change over time as vendors start building hype around it.
If marketing, finance and sales is dependent on a centralised data team for every new thing, the data team quickly becomes the bottleneck, stifling innovation and frustrating teams. Incorporating the principles of a Data Mesh enables those teams to manage their own data, according to well defined governance standards that enable interoperability.
The reality is that different teams are already managing their own data (via excel spreadsheets, web-apps, etc). If we can apply a bit more rigor to how these datasets are managed (e.g. so they can be shared, integrated, secured, etc), then the whole organisation benefits.
If Google were to kill it, you could easily run it on any other hosted Kubernetes service.
I haven't used Cloud AI Platform Pipelines, but have spent a lot of time working with Kubeflow Pipelines and its pretty great!
[1] https://github.com/kubeflow/pipelines
[2] https://www.kubeflow.org/docs/aws/ (Deploy to AWS)
[3] https://www.kubeflow.org/docs/azure/ (Deploy to Azure)
I love the product, but hate that I can't rely on it.
[0]: https://en.wikipedia.org/wiki/Thomas_Kuhn
[1]: https://en.wikipedia.org/wiki/The_Structure_of_Scientific_Re...
I'd like to look into it further, but I can't find any information about licensing. Given that there is no source code in the repo, it appears that this isn't an open source project.
Is it possible to view / check / correct any information about myself if I am a "passive user" on your platform?
Many eu countries and Austalia have privacy laws around these basic rights.
We have our own retry logic, which also logs the issue so we are aware of how frequently errors occur while a command / transaction is being executed.
This is using SQL Azure with the "Business" tier, so it will be interesting to see how the new (much more highly priced tiers) Standard and Premium tiers go.
Azure is a bit hit and miss. Its brilliant for getting something up an running quickly (using websites / SQL Server), but is a little flaky at scale.
Key problems include connection issues with SQL Server, connection issues with their hosted Redis service, pricing of SQL Server when using advanced features like geo-replication.
All in all though, its a pretty good development experience once you get your head around the fact that in the cloud services fail and there is nothing you can do about it except plan for it.
Oh and the Bizspark program they have gives you $100 worth of free hosting on Azure which is always nice.
Coming from that background, C and especially C# must seem extremely verbose.
For example (from Wikipedia):
In K, finding the prime numbers from 1 to R is done with [0]:
(!R)@&{&/x!/:2_!x}'!R
And APL[1]: (~R∊R∘.×R)/R←1↓ιR
Its truly awesome stuff.[0]: http://en.wikipedia.org/wiki/K_(programming_language) [1]: http://en.wikipedia.org/wiki/APL_(programming_language)
The dirty little secret of the business intelligence / dashboard industry is that no one logs into them.
A daily email helps with this problem, as people tend to read emails, even if its only a glance.
The issue is that composability is often tied to actually moving data around in the database which has terrible performance. That is, you can compose a query of multiple queries that dump partial data sets into temp tables.
Views get you part of the way there, but they are designed to be long lived and are visible to all database users until they are dropped. This means its dangerous to change them or clean them up, as its not always clear who they are being used by.
Ephermal or temporary views that are session/connection based, or even loadable as modules would be useful to me.
I would prefer a syntax layer that can be compiled / transformed back to SQL but that does basic things like having a query start with the tables, then joins, then groupings then the final projection.
Also a less cumbersome way to use the "WITH" statement to form named sub-queries.
Perhaps something like:
SELECT
COUNT(*) as columns,
column_type,
table_name
FROM (
SELECT c.id,
c.type AS column_type,
t.name AS table_name
FROM tables t
INNER JOIN columns c
ON t.id = c.table_id
WHERE t.system=false;
) a
HAVING COUNT(*) > 1
ORDER BY columns DESC
Being re-written as: # Use ":=" to replace WITH for named ephermal views
# Replace "WHERE" with "?", "SELECT" with "|>" at the end
non_system := tables
? system=false
|> name:table_name, is:table_id
# Replace INNER JOIN with "*="
non_system_columns := non_system.table_id *= columns.table_id
|> c.id, c.type:column_type
# GROUP BY columns are automatically generated by non-aggregated columns
column_types := non_system_columns
|> COUNT(*):columns DESC, column_type, table_name
So the final query may look like: non_system := tables ? system=false
|> name:table_name, is:table_id
non_system_columns := non_system.table_id *= columns.table_id
|> c.id, c.type:column_type
non_system_columns
|> COUNT(*):columns DESC, column_type, table_name
Any thoughts on this?This means that your data (in files) needs to be in the working directory, and is versioned along side your code. Sounds pretty cool, but I am not sure how it would scale for large / constantly changing data sets.
Unfortunately my carpentry skills lacking, and the Dell Single Monitor Arm [1] are NOT suitable for mounting 2 monitors on top of each other (not enough height).
[1]: http://accessories.us.dell.com/sna/productdetail.aspx?c=us&l...
The article was just supposed to be a bit of fun about my personal experience with different monitor configurations. I thought that this might be interesting to the HN crowd since so many of us spend so much time in front of screens.
Also, HN seems to have dropped the "∝" symbol from the title!
Essentially you rank your skills based around enjoyment and experience and it tells you what types of roles might be of interest to you.
This seems to be due to how much memory is accessed when checking a single bit, and the difficulty in predicting branches.
Of course it could just be that all my implementations have sucked, but even in playing around with libcds[1] didn't yield the kind of performance expected.
If anyone knows of a fast implementation of Wavelet Trees, I'd love to see it.
E.g. would it be possible to run a query against Red Shift (step 1), then cleanse it in Python (step 2), run R scripts over the Python output (step 3) then dump the results back to Red Shift (step 4)?
If I then decide I need to change the Red Shift query (step 1), can I then re-run the whole pipeline?
Munging data between different tools is what I seem to spend most of my time doing, so anything that helps that would be a big productivity boost.
There is even the phenomenon known as Chromothripsis [1], which has the entire genome smashed into thousands or millions of pieces and then randomly put back together.
Micro-array analysis (as opposed to the Exome sequencing they are starting to offer, which may be useful for BRCA diagnosis) focusses on frequently occurring single nucleotide polymorphisms (substitution of one base for another). The majority of these are benign.
By looking at lots of commonly occurring variants it is hoped that they can map regions of the genome that may harbour other mutations that may actually have an effect.
Baynes et al., 2007 [1] showed that none of the commonly occurring SNP's in BRCA (> 1-5% of the population) have a significant risk in cancer.
When I first started with golang I too was flailing around with deadlocks, but once you have spent a bit of time with it, it becomes obvious when you have messed up.
I think that this is more of a lack of familiarity with the language rather than an intrinsic problem with golang.
"This new threshold applies only to petitions created from this point forward and is not retroactively applied to ones that already exist."
So the Ortiz petition will be unaffected.
Or does the DoJ have to get approval from victims before offering a plea deal? I can't imagine the victim of an assault or rape being terribly happy "signing off" on a plea for their attacker.
From a bit of reading it appears that victims have to be informed of any plea bargain and that their needs must be "taken into consideration", but I cannot find anything that states that the plea must be "signed off" by the victim.
If this is the case, then this story seems rather flawed.