785 karma · joined September 13, 2009
http://ogrisel.com http://twitter.com/ogrisel http://github.com/ogrisel
Both papers are direct applications of the chain rule applied to estimate the gradient of a multivariate function.
We need local sandboxing for FS and network access (e.g. via `cgroups` or similar for non-linux OSes) to run these kinds of tools more safely.
This is why it's important to describe explicitly the three points in text:
- steps to reproduce;
- what you expected to happen;
- what actual result you observe instead.
Something that might be obvious to you but isn't for others will just be silently ignored most of the time.
EDIT: I now see the problem after reading your other reply above:
https://news.ycombinator.com/item?id=46064757#46069546
This is why it's important to describe explicitly the difference between what you expected and what you observed. I swear I did not see the change in button width before reading the linked comment.
- steps to reproduce from scratch;
- what you expected to happen;
- what you actually observed (include the screenshot or video capture in addition to a textual description).
Otherwise, you might risk your report being ignored due to a silent misunderstanding about the mismatch between your expectations and the actual results.
If you want to share structured Python objects between instances, you have to pay the cost of `pickle.dump/pickle.dump` (CPU overhead for interprocess communication) + the memory cost of replicated objects in the processes.
https://github.com/allenai/OLMoE
Algorithmic puzzles, on the other hand, both require reasoning and are easy to verify.
There are other things in coding that are both useful and easy to verify: checking that the generated code follows formatting standards or generating outputs with a specific data schema and so on.
- open-source inference code
- open weights (for inference and fine-tuning)
- open pretraining recipe (code + data)
- open fine-tuning recipe (code + data)
Very few entities publish the later two items (https://huggingface.co/blog/smollm and https://allenai.org/olmo come to mind). Arguably, publishing curated large scale pretraining data is very costly but publishing code to automatically curate pretraining data from uncurated sources is already very valuable.
The fact that DeepSeek-R1 is so much better than DeepSeek-V3 at various important tasks means that Chain-of-though / thinking-before-answering models are better. But they are also more compute intensive at inference time than their instruction non-thinking counterparts.
So even if the DeepSeek-V3 pretraining + GRPO COT post-training procedure was cheaper than anticipated to reach o1 grade performance, inference is still costly, even if you use a distilled model.
This way, pip will fail if a dependency does not provide a `.whl` package, instead of automatically falling back to the "build from source" mode that can lead to arbitrary code execution at install time (via setuptools' `setup.py` or any other build backend mechanism).
However, installing from wheels just protects from arbitrary code execution at install time. If you do not trust the source and integrity of the package you install, you would still be subject to arbitrary code execution at import time.
Therefore, tools and processes to improve package provenance tracing and integrity checking are useful for both kinds of installations.
- https://github.com/scipy/scipy/issues/21479
it might be the same as this one that further involves OpenMP code generated by Cython:
- https://github.com/scikit-learn/scikit-learn/issues/30151
We haven't managed to write minimal reproducers for either of those but as you can observe, those race conditions can only be triggered when composing many independently developed components.
Sometimes the optimal level of parallelism lies in an outer loop written in Python instead of just relying on the parallelism opportunities of the inner calls written using hardware specific native libraries. Free-threading Python makes it possible to choose which level of parallelism is best for a given workload without having to rewrite everything in a low-level programming language.
https://data-apis.org/array-api/
So it's possible to write array API code that consumes arrays from any of those libraries and delegate computation to them without having to explicitly import any of them in your source code.
The only limitation for now is that PyTorch (and to some lower extent cupy as well) array API compliance is still incomplete and in practice one needs to go through this compatibility layer (hopefully temporarily):
If the notebooks themselves contain assertions to check that expectations on the outputs are met, then you have an automated way to check that the notebooks behave the way you want on some test inputs. For long notebooks, this is more like integration/functional tests rather than unit tests, but I think this is already an improvement over manually run notebooks.
Note sure about strict types: you mean running mypy on a notebook? Maybe this can be helpful:
- https://pypi.org/project/nb-mypy/
About linters, you can install `jupyterlab-lsp` and `python-lsp-ruff` together for instance.
I don't know the average energy/hardware*time usage per query on google translate vs competing LLMs such as Claude 3 Opus but I wouldn't be surprised that a large LLM such as Claude 3 Opus would be much too expensive to be used as the backend model for a free service like Google Translate.
The paper authors do acknowledge this concern and run experiments on smaller models with knowledge distillation. However, as far as I know we cannot know if their distilled networks can compete with the current Google Translate system in terms of energy / hardware usage efficiency.
https://www.maersk.com/news/articles/2023/06/13/maersk-secur...
The cleanest process combines hydrogen produced from renewable electricity (typically wind or solar) and CO2 captured from the combustion of biomass.
But if there non-trivial logic in the code of the tests, I agree this is probably a risky approach.
- they require different instructions for each platform: one for each distribution: it's not possible to give a one liner that will work for all supported OS/architecture combinations, - distributions are not all updated at the speed of upstream. Sometimes users do not want to wait several years to use a recently released feature.
- https://peps.python.org/pep-0574/
and also provides extra API to handle large data buffers externally ("out-of-band") via custom callbacks. This is in addition to the no-copy semantics memory optimization when loading/storing such arrays "in-band" without providing custom callbacks.
import pickle
with open("model.pkl", mode="wb") as f:
pickle.dump(trained_model, f, protocol=pickle.HIGHEST_PROTOCOL)
with open("model.pkl", mode="rb") as f:
trained_model = pickle.load(f)
pickle from the standard library with protocol 5 can store and load large data buffers often found as attributes of scikit-learn models (typically large numpy arrays) without extra memory copies (as joblib.dump and joblib.load were designed to do with a few hacks that violate the official pickle protocol).A new safer alternative for scikit-learn model persistence is skops:
- https://skops.readthedocs.io/en/stable/persistence.html
It makes it possible to trust a list of types of Python objects that are safe to load and refuse to load skops files with untrusted types.