I hope, since it’s opensource, you are thinking about exposing api / hooks for downstream tasks.
35 karma · joined October 11, 2018
I hope, since it’s opensource, you are thinking about exposing api / hooks for downstream tasks.
Speaking of implementation, i don’t mind if a browser extension forward cookies from my browser to the automation (privacy and security is an issue of course, and i’d ideally want the cookies to not leave my device, but personally i’m okay with some trade off).
Our main requirement was querying recent operational data across daily/weekly/monthly/quarterly timeframes. The data sources included OLTP binlogs, OLAP views, SFDC, and about 15 other marketing platforms. We implemented a datalake with our own query and archival layers. This approach worked well for queries like "conversion rate per channel this quarter" where we needed broad data coverage (all 17 integrations) but manageable depth (reasonable row scanned).
This architecture also enabled quick solutions for additional use cases, like on-the-fly SFDC data enrichment that our analytics team could handle independently. Later, I learned the team integrated LLMs as they began dumping OLAP views inside the datalake for different query types, and eventually replaced our original query layer with DuckDB.
I believe approaches like these (what I had done as in house solution and what definite may be doing more extensively) are data and query-pattern focused first. While it might initially seem like overkill, this approach can withstand organizational complexity challenges - with LLMs serving primarily as an interpretation layer. From skimming the Dustt blog, their approach is refreshing, though it seems their product was built primarily for LLM integration rather than focusing first on data management and scale. They likely have internal mechanisms to handle various use cases that weren't detailed in the blog.
- if the key is present
- if the value of the key is in given range (any filter function you can imagine)
- conditional key presence (Ex: if key1 is present key2 should be less than 10)
- complex nested structure of the payload
- array validation (length / type / property)
I end up write a significant amount of if else block just to validate if the payload is acceptable. So I wrote a tool a long time back which I use in all my JS based server project (My gateway servers are always in Node for past 3-4 years because of the amount of instrumentation / ease of debugging)
I wrote helson 2 years back, where a dev would define schema of a payload
``` typedef Payload { str "url": pass, []str "tags": pass, bool "isCollection": pass, obj "content": { str "body": strShouldNotBeEmpty | shouldBeOfMinimumLength 15 | shouldHaveMinimumWordLength 5, }, } ```
with primitive type support of str, bool, number and compound type support array, enum, ref (https://github.com/adotg/helson/blob/develop/test/helson.tes...) for type checking and custom function (strShouldNotBeEmpty, shouldBeOfMinimumLength above example) as value checking
And it tests against incoming payload
``` { url: '#/e/hash-of-a-link', tags: ['a1', 'a2', 'a3'], content: { body: 'This is a body', } ```
Now that I use typescript, I was in the verge of deciding (I don't have enough time solving just for myself as the value addition is not justified for myself, but if enough people wants it I'll build it) should I also build a typescript supported workflow. Like from a schema like above generate interfaces (and vice versa). Or is this lib meaningful to you at all.
It also throws meaningful error automatically (which you can override) to return to client directly
I have noticed people who are extra organized in real life, who keeps every single file in a right directories after download tend to have inclination for prematured code refactoring for future use.
If these guys become code architect then i end up doing so many unnecessary things. The fundamental assumption of the refactoring gets changed very fast and the code needs to be rewritten for the majority of the cases.
From a company level I find the instructions are clear, mostly holding the same contract between services as long as possible and less schema change.
Its the engineer with subjective idea of perfect code / supporting future work makes it even more complex
Do you have any suggestions for testing (Mostly service API testing) ?
Which is why performing this in browser env even for low amount of data (say 10k) is nightmare. There are ways you can address this but while in browser you hit the limit pretty soon.
We wanted the concept to be validated first hence we have build it for browser only. But would love to hear / learn / discuss with you on this before we go ahead and build the data model in server.
Another ambiguity with the interaction is visual effect of interaction. Questions like do you really want all your chart to be cross connected. A in house survey showed us there is no certainty of the answer. And what kind of visual effect should happen on interaction differs person to person and is a function of use case. Which is why we have chosen go for chosen behaviour like
``` muze.ActionModel.for(...canvases) / for all the chart in page / .enableCrossInteractivity() / allow default cross interactivity / .for(tweetsByDay, tweetsByDate) / but for first two canvas in the example / .registerPropagationBehaviourMap({ select: 'filter', brush: 'filter' }) / if selection using mouse click or brushing happens filter data / ```
we are still writing docs for this. We hope to finish all the these docs in two weeks time.
https://jsfiddle.net/adarshlilha/8n5a94j1/20/
Does this help?
The web framework fetches data, does some additional checking on data and schema, process visible code and render it. That is probably the reason you are seeing lag.
Also there are few areas where Muze performance needs to be improved. We are doing a release to address this soon.
- you feel you had hard time achieving with Plotly? - a feature you wanted is not supported by Plotly?
So here is the thing with our DataModel. Every time you perform an ops on DataModel it create another instance. Now performing multiple such actions create a DAG where each node is an instance of DataModel and each edge is an operation.
We have auto interactivity, which propagates data (dimensions) pulse along the network. Any node which is attached to visualiztion receives those pulses and changes the visual.
So far I have not found any relational interface which exposes this DAG graph and api to user. Hence we though of building this.
Having said that, we might use some established relational interface and do the propagation ourself.
WebAssembly is on our radar and is coming soon. But we just wanted to release a super early version of what we build so far.
Will figure out the effort and roadmap and then keep you updated on the plan.
Will upload the docs for this soon.
Vega-Lite paper and layered grammar of graphics are the biggest motivations to write Muze. Vega-Lite is still my goto viz library for my ipython and JS work. Hence there are healthy intersections between vega-lite and muze terminologies and concepts.
Would definitely love to check new development.
For the time being, you can click on the play button on top right corner of the code section.
However, will make sure it gets tested with all linux dist before next release.
Are you looking for integration? > Also what is the Reactjs story here ?
However https://www.charts.com/muze/docs/composing-layers explains some part about muze's composability. And this is an example https://www.charts.com/muze/examples/view/composition-of-lay...
Apart from composable layers Muze has - tabular layout (visual crosstab) created from data facets https://www.charts.com/muze/examples/view/crosstab-chart - auto interactions https://www.charts.com/muze/examples/view/crossfiltering-wit... - Legend on any chart https://www.charts.com/muze/examples/view/gradient-legend
etc...
We would love to know your - use case - number of data points - ops on data on serverside
You can mail us to eng@charts.com
However, every operation does not support formula storing. Operation like joining, grouping creates new data.
We are updating the docs rapidly. All this info would be on the docs soon.
Vega is descriptive version of d3. We find it hard for debugging and creating complex viz.
Vega-lite is concise and more intuitive version of vega though.
However Muze was created to start directly from data, creating layout, composable layers, automatic cross interaction and a robust interaction mental model. Muze is inspired from VegaLite-InfoVis and Snap together viz paper.