It would be nice if you could post if the actual data matches your reconstruction—now that you have it in hand. Would help us not worry about the data provenance and focus on the result you found.
582 karma · joined October 13, 2014
Founder, Synthetic Minds YC S'18 - Program synthesis for desktop automation https://warpdrive.co
Founder, 20n YC W'15 - Program synthesis for synthetic biology: http://20n.com
Postdoc (UC Berkeley): Program synthesis for cell designs to engineer cells using synthetic biology.
PhD (University of Maryland): Program synthesis.
Personal page: http://www.saurabh-srivastava.com
It would be nice if you could post if the actual data matches your reconstruction—now that you have it in hand. Would help us not worry about the data provenance and focus on the result you found.
* qwen-{1.8B,7B,14B}:
* 3 trillion tokens; start with BPE tiktoken, cl100k base vocab, augmented with chinese, numbers split into digits, final vocab 152k.
* RoPE - rotary positional embedding
* context length 2048
* qwen-14b perf percentages: 66.3 MMLU(5), 72.1 CEval(5), 61.8 GSM8K(8), 24.8 MATH(4), 32.3 HumanEval(0), 40.8 MBPP(3), 53.4 BBH(3); beats LLaMA2-13B on all, but behind LLaMA2-70B on all except CEval, MATH and HumanEval (somewhat surprising)
* code-qwen-{7B,14B} * additional 90B code tokens over base
* context length 8192, flash attention
* 14B perf: humaneval 66.4, mbpp 52.4; ok, but not stellar (similar numbers as OSS wizardcoder-py, and lower than gpt-3.5)
* math-qwn-{7B,14B}-chat * math instructional dataset
* context length 1024
* 14B perf: gsm8k 69.8, MATH 24.2, Math401 85.0, Math23K 78.4 (substantially better than OSS in the same weight class (WizardMath and GAIRMath-Abel) on MATH but same ballpark on GSM8k -- surprising). Math23K is chinese grade school math; and Math401 is arithmetic ability.
* comprehensive automatic evaluation in Appendix A.2.1 pg 36 (based on OpenCompass'23)* chat format: <|im_start|>system You are a helpful assistant.<|im_end|> <|im_start|>user Hello!<|im_end|> <|im_start|>assistant Hello! How can I assist you today?<|im_end|>
LLM frameworks might imply stack for building models (pertaining, fine tuning, inference etc)
* In standard attention in transformers, cost scales quadratically with length of sequence, which restricts model context. This work presents subquadratic exact operator allowing it to scale to larger contexts (100k+).
* They introduce an operator called "Hyena hierarchy", a recurrence over 2 subquadratic operations: long convolution, and element-wise mul gating. Sec 3.1-3.3 define the recurrences, matrices, and filters. This is importantly, a drop in replacement for attention.
* Longer context: 100x speedup over FlashAttention at 64k context (if we view flash attention as an non-approx engg optimization, then this work is improving algorithmically, and getting OOM over that). Associate recall, i.e., just pull data, show improvements: Experiments on 137k context, and vocab sizes of 10-40 (unsure why they have bad recall on small length sequence with larger vocab, but they still outperform others)
* Comparisons (on relatively small models, but hoping to show pattern) with RWKV (attention-free model, trained on 332B tokens), GPTNeo (trained on 300B tokens), with Hyena trained on 137B tokens. Models are 125M-355M sized. (Section 4.3)
* On SuperGLUE, zero-shot and 3-shot accuracy is ballpark similar to GPTNeo (although technically they underperform a bit for zero-shot and overperform a bit for 3-shot). (Table 4.5 and 4.6)
* Because they can support large (e.g., 100k+) context, they can do image classification. They report ballpark comparable against others. (Table 4.7)
Might have misread some takeaways; happy to be corrected.
A few questions that might help an enterprise customer: How big is your base model? Where did you find more datasets (maybe just a hint would be sufficient)? Are you using SantaCoder [3]? Anything you can say about your fine-tuning that makes it special? Totally on board with you that HumanEval/MBPP are not great benchmarks for real world, and do you have a suggested alternative to help me see the value?
The calculus for an enterprise customer might be: "We could fine tune a 6B model on our internal code and internal benchmarks (say with a month of work, a few thousand in compute, 2 people on task), but I'd rather buy an off-the-shelf solution like codecomplete.ai. They give us XYZ benefits." Articulate the XYZ for a technical decision maker who will be your target audience.
* [1] https://huggingface.co/datasets/bigcode/the-stack
* All variants were trained on 1T - 1.4T tokens; which is a good compared to their sizes based on the Chinchilla-metric. Code is 4.5% of the training data (similar to others). [Table 2]
* They note the GPU hours as 82,432 (7B model) to 1,022,362 (65B model). [Table 15] GPU hour rates will vary, but let's give a range of $1 to $4. The 7B model would have cost ~$82-329k and the 65B something in the range of ~$1-4M. They also note their total time spent for all models: "we used 2048 A100-80GB for a period of approximately 5 months" [sec 6, pg 10]
* 65B model's performance is broadly comparable to PALM-540B. Not a small feat, but also could indicate the benefits of good model-vs-token size ratios [Tables 3,4,5,6]. Their conjecture for underperforming on MMLU (multitask language understanding) compared to PALM-540B and Chinchilla-70B is smaller fraction of books and academic training data.
* Math and code tasks: Math tasks they are substantially worse than Minerva (comparing their 65B to Minerva 62B; they hands down fail against Minerva 540B) [Table 7]. Code tasks they are broadly competitive with PALM-540B (HumanEval and MBPP evals) [Table 8]
* Surprising that instruction fine tuning takes such a small part of the paper (sec 4, pg. 7)
* "We propose to extract memorized images by generating many times with the same prompt and flagging cases where many of the generations are the same."
* "- Diffusion models memorize more than GANs - Outlier images are memorized more - Existing privacy-preserving methods largely fail"
* "Stable Diffusion is small relative to its training set (2GB of weights and many TB of data). So, while memorization is rare by design, future (larger) diffusion models will memorize more."
* "It only memorizes a very small subset of the images that it trains on."
* "our goal is to show that models can output training images when generating in the same fashion that normal users do."
* Initialization was done 42 days ago: https://etherscan.io/tx/0x53fd92771d2084a9bf39a6477015ef53b7... -- "Click to see More" and notice "Input Data" parameter [2] which sets _committedRoot to 0x00.
* Click through the To contract to get to the code (click on Contract tab): https://etherscan.io/address/0xb92336759618f55bd0f8313bd8436...
Just adding direct links to what samczsun and 0xfoobar are talking about in https://twitter.com/samczsun/status/1554260106107179010 and https://twitter.com/0xfoobar/status/1554269071214088193/phot...
This should almost always get you the answer in 3 guesses. [Edit: Realized after writing that 3 is too optimistic given you might not have enough constraints after the 2nd guess, and agree with dzdt's comment. 4 might be where we end up at.]
I would be curious to see the distribution of #guesses needed across the 158k words.
Not having to dig up COM UUIDs for interfaces is super nice. The lead developer of com-rs told me that com-rs is deprecated in favor of windows-rs (he works on both.)
Synthetic Minds is building program synthesizers, i.e., automation that can write code. We have a working prototype in stealth and are currently in the process of doing user studies.
We are looking for:
- Senior full stack / generalist
- UI/UX designer
Details on the company and open positions: https://www.workatastartup.com/companies/1873
The engineering team works all the way from the front-end to the bleeding-edge backend program synthesis stack. You will learn a lot, as I am sure we'll learn from you. The team is qualified to build this: there are 3 PhDs, and another 3 with 10+ yrs of product/engineering/frontend experience.
We are backed by YC, Khosla, and Pantera. This is my 2nd YC startup. The current team of 8 is spread across the Bay Area, Seattle, and Utah. Remote positions in the US work for us.
Contact me at saurabhs@synthetic-minds.com
[1] https://www.daxx.com/blog/development-trends/number-software...
Synthetic Minds is building program synthesizers, i.e., automation that can write code. We have a working prototype in stealth and are currently in the process of doing user studies.
Our hiring needs over the next month are:
- Full stack / frontend engineer
- Generalist
- Passively looking: UI/UX designer
Details on positions we are actively hiring for are on: https://www.workatastartup.com/companies/1873
The engineering team works all the way from the front end to the bleeding-edge backend program synthesis stack, so there is opportunity for significant technical learning. Our team includes 3 PhDs and 2 ex-Googlers, and people with 10+ years of experience. We are backed by YC, Khosla, and Pantera. This is my 2nd YC startup. Our team of 6 is split across SF + Seattle, and if you are remote we'd prefer the US to maximize time-zone overlap.
Contact me at saurabhs@synthetic-minds.com
Synthetic Minds is building program synthesizers, i.e., automation that can write code. We have a working prototype in stealth and are currently in the process of doing user studies.
Our hiring needs over the next month are:
- Full stack / frontend engineer
- Generalist
- Passively looking: UI/UX designer and Product Manager
Details on positions we are actively hiring for are on: https://www.workatastartup.com/companies/1873
All these people will get exposed to a bleeding-edge program synthesis stack, so there is opportunity for significant technical learning. We are an all engineering team (including 3 PhDs and 2 ex-Googlers) backed by YC, Khosla, and Pantera. This is my 2nd YC startup. Our team of 5 is in Seattle + SF, and if you are remote, we'd prefer the US to maximize time-zone overlap.
Contact me at saurabhs@synthetic-minds.com
Scanning my own face, for each of these 3, i tried to see my own saccade, and (1) I am blind to it in the mirror; (2) barely see it on the iphone, and (3) very visibly see it on the macbook.
Synthetic Minds is building program synthesizers, i.e., automation that can write code. We have a working prototype in stealth and are currently in the process of doing user studies.
Our hiring needs over the next month are:
- Full stack / frontend engineer
- UI/UX designer
- Generalist that can go across the stack
All these people will get exposed to a bleeding-edge program synthesis stack, so there is opportunity for significant technical learning. We are an all engineering team (including 3 PhDs and 2 ex-Googlers) backed by YC, Khosla, and Pantera. This is my 2nd YC startup. Our team of 5 is in Seattle + SF, and if you are remote, we'd prefer the US to maximize time-zone overlap.
Contact me at saurabhs@synthetic-minds.com
Synthetic Minds is building program synthesizers, i.e., automation that can write code. Think of what we are building as a compiler that takes code and translates it to theorem proving, so that we can build automation that can understand code almost as close to a human. If it can understand code, with sufficient compute it can even synthesize it.
For the kinds of technical problems we handle, look at https://synthetic-minds.com/pages/conference/2019/#program
We are an all engineering team (including PhDs and ex-Googlers) backed by Y Combinator, Khosla Ventures and Pantera Capital. We are currently in Seattle and SF and are looking for people with significant engineering experience.
Contact me at saurabhs@synthetic-minds.com
Synthetic Minds builds program synthesizers, i.e., automation that can write code. There is two decades of research that forms the backbone of this tech. The founder has a PhD in the domain, and the CTO is an ACM Fellow with 20+ years of work in Program Synthesis (https://synthetic-minds.com/pages/jobs.html#about). We have raised $5.6M from YC, Khosla Ventures, and Pantera Capital.
We are an all engineering team, and are looking for engineer #7, ideally with a masters or PhD (or built a relevant well-known project.) Programming languages, compilers, formal methods, SMT solving (Z3) are relevant topics for us. For the kind of work you'll be doing, see the technical program here: https://synthetic-minds.com/pages/conference/2019/.
Email saurabhs@synthetic-minds.com for more details.
Synthesizing smart contracts from test cases. https://synthetic-minds.com/pages/blog/blog-2019-09-12.html
E.g. SMT solvers (CVC, Z3, or the like) -- infinitely more fun; and require experience to truly understand what works and what doesn't. Or if you've done something really novel with meta-programming or designed a custom DSL for a domain.
The researchers will talk about their peer-reviewed work in web automation, hardware security, operating system extensions, programming for non-programmers, automatic code translation, and superoptimization. Hopefully, this will illustrate the power and limitations. We'd love for people to extrapolate from these onto their own domain-specific automation needs.
By touching upon foundational techniques (making imperative code functional, symbolic compilation, SMT encodings, partial evaluation), hopefully the leap to "code synthesis" will seem less like magic and more like an obvious next step. In addition, open-source frameworks exist (e.g., Rosette, Sketch) that abstract away these foundations, and the program will cover those in hands-on workshops.
We’d love to hear insights into application from people for whom synthesis is new. Some problems are exciting to us (Synthetic Minds is working on smart contract synthesis); and we’d love for the community to brainstorm applications to their domains.
The ideal candidate has a master/phd in systems, compilers, programming languages, or distributed systems. Synthetic Minds will allow you to leverage your technical chops.
Synthetic Minds is building program synthesizers, i.e., automation that can write code. We have a system in production that reads/writes smart contracts in Ethereum's Solidity language, and we use it to ensure our customer's code is secure and correct. Eventually, we plan on going far beyond smart contracts. Think of what we are building as a compiler that takes code and translates it to theorem proving, so that we can build automation that can understand code almost as close to a human. If it can understand code, with sufficient compute it can even synthesize it.
In Oct 2018, we raised a $5.6M seed round from Y Combinator, Khosla Ventures and Pantera Capital [3]. We have paying customers and a backlog waiting to be on-boarded. This is the founder’s second startup and they have a PhD in the area. The 1st employee was the first hire at Parse (YC S11) and has 10 yrs at Google. We are currently hiring engineer #5, and aim to be a 15 person all-engineering team in 2019.
Roles/Openings [1]:
* Systems/infrastructure engineer — You’ll be working on distributing heavy CPU processes on AWS. Making sure processes run reliably over many days. Ensure robustness of the infrastructure across node/process/memory/algorithm failures.
* Compilers/verification/synthesis engineer — You’ll be working on developing new algorithms that analyze and generate code [2]. You’ll identify when an engineering solution is needed (i.e., throw across a cluster of machines), or when an algorithmic improvement is required. You might even play with the Z3 theorem prover. And if you’re really into it, you can improve Z3.
* Smart contract engineer — You'll be working on the front-end of the compiler, which reads in smart contracts languages (e.g., Solidity) and makes it accessible to the backend (the part that does semantic analysis).
Contact: saurabhs@synthetic-minds.com - Saurabh, Founder
[1] Synthetic Minds Jobs: https://synthetic-minds.com/pages/jobs.html
[2] Program synthesis: https://medium.com/@vidiborskiy/software-writes-software-pro...
[3] Forbes funding article: https://www.forbes.com/sites/darrynpollock/2018/10/22/invest...
We are not front-end people :) and it wasn't an active choice. Let me know if you see anything else. Sincerely appreciate the help.
The ideal candidate has a master/phd in systems, compilers, programming languages, or distributed systems. Synthetic Minds will allow you to leverage your technical chops.
Synthetic Minds is building program synthesizers, i.e., automation that can write code. We have a system in production that reads/writes smart contracts in Ethereum's Solidity language, and we use it to ensure our customer's code is secure and correct. Eventually, we plan on going far beyond smart contracts. Think of what we are building as a compiler that takes code and translates it to theorem proving, so that we can build automation that can understand code almost as close to a human. If it can understand code, with sufficient compute it can “synthesize” it.
In Oct 2018, we raised $5.6M from Y Combinator, Khosla Ventures and Pantera Capital [6]. We have a backlog of customers waiting to be on-boarded. The team is experienced. This is my 2nd YC startup and I have a PhD in Program Synthesis. The 1st employee was the first hire at Parse (YC S11) and spent 10 yrs at Google. We aim to be a 10-15 person all-engineering team in 2019.
Roles/Openings (see [1]):
# Software engineer: Systems/infrastructure — You’ll be working on distributing heavy CPU processes on AWS. Making sure processes run reliably over many days. Ensure robustness of the infrastructure across node/process/memory/algorithm failures.
# Software engineer: Compilers/verification/synthesis — You’ll be working on developing new algorithms that analyze and generate code [2]. You’ll identify when an engineering solution is needed (i.e., throw across a cluster of machines), or when an algorithmic improvement is required. You might even play with the Z3 theorem prover [3]. And if you’re really into it, you can improve Z3.
# Software engineer: Smart contracts — You’ll be working on the “front-end of the compiler”, which reads in smart contracts languages (e.g., Solidity) and makes it accessible to the backend (the part that does semantic analysis). Desire to work at the compiler level of smart contracts is required, e.g., see [4] — experience in writing smart contracts is easily acquired as a side effect.
Contact: saurabhs@synthetic-minds.com - Saurabh, Founder
[1] Synthetic Minds Jobs: https://synthetic-minds.com/pages/jobs.html
[2] Program synthesis: https://medium.com/@vidiborskiy/software-writes-software-pro...
[3] Z3 Theorem Prover: https://github.com/Z3Prover/z3
[4] Solidity AST: https://github.com/ethereum/solidity/blob/develop/libsolidit...
[5] Solidity smart contracts: https://solidity.readthedocs.io/en/v0.4.21/solidity-by-examp...
[6] Forbes funding article: https://www.forbes.com/sites/darrynpollock/2018/10/22/invest...