105 karma · joined December 12, 2016
Presently, a large language model (LLM) uses the same amount of computing resources for both simple and complex problems, which is seen as a drawback. Imagine if an LLM could adjust its computational effort based on the complexity of the task. During inference, it might then perform a sort of search across the solution space. The "search" mentioned in the article means just that, a method of dynamically managing computational resources at the time of testing, allowing for exploration of the solution space before beginning to "predict the next token."
At OpenAI Noam Brown is working on this, giving AI the ability to "ponder" (or "search"), see his twitter post: https://x.com/polynoamial/status/1676971503261454340
""" Caching Libraries
joblib.Memory provides caching functions and works by explicitly saving the inputs and outputs to files. It is designed to work with non-hashable and potentially large input and output data types such as numpy arrays.
"""
From https://pypi.org/project/diskcache/https://github.com/stanfordnlp/dspy/blob/main/dsp/modules/ca...
Yet, that still means, knowledge of O cannot be synthesized.
Each of the multiplexed channels are individually limited by the Shannon limit, and with higher power the fiber's Kerr effect creates interference which creates a sweet spot for the optimal optical launch power.
the novelty here is that the spectral channels are all generated from a single laser source rather than a laser per channel
In an experiment like this, only the initial light source is modulated and therefore all channels carry the same data. The equipment for the transmitter and receiver chain is so expensive that university labs can barely afford one of each.
Not sure what the baud rate of a single channel was in their experiment but probably between 32-80Gb which is common for the lab equipment at Universities. The industry is knocking on 100-400Gb where for the actual decoding and signal processing there is massive parallelism applied to reduce the rate even more
Description: https://danishdesignaward.com/en/arkiver/nominee/adam-lenzin...
https://www.nts.live/shows/the-do-you-breakfast-show/episode...
https://video.ethz.ch/speakers/bernays/2019/7b11b50e-f813-4d...
I loved this book: https://users.aalto.fi/~ssarkka/pub/cup_book_online_20131111...
and also Thomas Schoen group does great work on Sequential Monte Carlo (SMC), MCMC for sequential data :) http://user.it.uu.se/~thosc112/index.html
They are also building a probabilistic programming language for sequential data! https://github.com/lawmurray/Birch
Salary-wise it's difficult for me to answer, as most of the big industry players are in the US or Canada, but I'm in Europe. From hearsay a fresh PhD with reasonable publication list will get around 10k per month in the bay area. However, glassdoor might give you a better idea. Look for companies like: Infinera/ciena/acacia communications/juniper/mellanox/finisar/keysight...
After my PhD, I probably could have pursued a postdoc somewhere in Europe, but I decided to leave academia. I wanted to stay in Copenhagen and therefore had to change my field of work and will be working for a hearing aid company in their signal processing department.
1) The fiber is pretty thin, but there are several fibers combined in one transoceanic cable, see here https://en.wikipedia.org/wiki/Submarine_communications_cable. They don't patch them, but splice the fibers. Patching introduces a little loss, so it should be avoided.
2) Not sure how deep they are buried, but I think to remember that the shore end of a submarine cable is better protected than the part in the deep ocean. Due to more ship traffic at the coast etc.
3) Good question. The loss of an optical fiber is roughly 0.2 dB/km. Across the ocean the optical signal must be amplified several times (every 80-100km). Nowadays the amplification is all optical (EDFA or Raman). Before there were electrical regeneration schemes. Check out the history of the field, it's is quite interesting [1].
4) Not sure how much data such a cable can carry, since it depends on how many fibers are deployed within. However, there are multiple interesting things to look into here. In research labs, people are investigating multicore/multimode fibers (space division multiplexing) [2], these fiber have incredible capacity. Personally, I think the most interesting metric is the spectral efficiency, so how much data can be transmitted per second per Herz. Such a metric is independent of multiplexing schemes over space/wavelength/time, and improvements have to come from better devices, signal processing or signal shaping methods [3]. Another mind-blowing area is using the nonlinear Fourier transform for better signaling methods [4].
Feel free to ask more questions :)
and checkout the two biggest conferences for more in depth info.
https://www.ofcconference.org/en-us/home/
https://www.ecocexhibition.com/
[1] https://www.osapublishing.org/oe/abstract.cfm?URI=oe-26-18-2...
[2] Multicore: https://ieeexplore.ieee.org/abstract/document/7341685
Multimode: https://www.osapublishing.org/abstract.cfm?uri=ofc-2018-Th4C...
Multimode+core: https://ieeexplore.ieee.org/document/8535233
[3] https://arxiv.org/abs/1606.04073
[4] https://www.osapublishing.org/optica/abstract.cfm?uri=optica...
could you elaborate on that? pro research or against?
Although, the gains are marginal, I like the method a lot since it combines physics (nonlinear schroedinger equation derived fiber channel model), information theory (optimizing for mutual information) and machine learning. It wouldn't have been possible without the people who published the fiber channel model, my colleagues and in particular the colleagues who could help me in the lab.
The debugging was hell, there are so many dimensions where stuff can go wrong (besides the usual bugs): physical parameters with the wrong unit, the implementation of the fiber model in tensorflow, the machine learning parts with its training process. Plus the things that can go wrong in the lab.
There are still pieces where I'm not 100% sure, and would love to speak to someone with some background in autoencoders.
I've open sourced the autoencoder and the fiber channel model, checkout my github with the same username as here!