CLI tools hidden in the Python standard library
til.simonwillison.net
til.simonwillison.net
You give it a pattern for each token type, and a function to be called on each match, and you get back a list of processed tokens.
Importantly, it processes the list in one pass and ensures the matches are contiguous, where a naive `re.findall` with capture groups will ignore unmatched characters. You also get a reference to the running scanner, so you can record the location of the match for reporting errors.
import re
scanner = re.Scanner([
(r"[0-9]+", lambda scanner, token:("INTEGER", int(token))),
(r"[a-z_]+", lambda scanner, token:("IDENTIFIER", token)),
(r"[,.]+", lambda scanner, token:("PUNCTUATION", token)),
(r"\s+", None), # None == skip token.
])
results, remainder = scanner.scan("45 pigeons, 23 cows, 11 spiders.")
assert not remainder
print(results)
[('INTEGER', 45),
('IDENTIFIER', 'pigeons'),
('PUNCTUATION', ','),
('INTEGER', 23),
('IDENTIFIER', 'cows'),
('PUNCTUATION', ','),
('INTEGER', 11),
('IDENTIFIER', 'spiders'),
('PUNCTUATION', '.')]
[0]: https://stackoverflow.com/a/693818/252218[1]: https://en.wikipedia.org/wiki/Lexical_analysis#Tokenization
Bummer, it's a cool feature, but I don't feel safe relying on undocumented features.
They regret including most modules… it seems they regret making python altogether instead of sticking with C? :D
I also find it odd. Python would probably be a little known language without the huge batteries included by default. It's invaluable when you're working in a environment where you can't fully control what is installed, what is the case of most people at work. I believe this crusade endangers the language long-term.
"gl", "sgi", "fl", "sunaudiodev", "audioop", "stdwin", "rotor", "poly", "whatsound", "gopherlib" .. the list goes on.
Some of these were dropped with 1.6 (see "Obsolete Modules" at https://www.python.org/download/releases/1.6/ ). Some with 3.0 (see https://docs.python.org/3.10/whatsnew/2.6.html for then-upcoming removals.)
(1) Not letting little used modules that pose maintenance problems drive up the cost of maintaining Python, and
(2) Not forcing actively used, actively developed modules to be limited to the core language upgrade cadence (and not forcing users to upgrade the language to get upgrades to the modules.)
Ruby I think has a decent approach to this (particularly, one that deals with #2 better than just evicting things entirely from the standard distribution) with “Gemification” of the standard library, where most things that are moved out of the standard library aren’t moved out of the standard distribution, but into a set of gems distributed with the standard distribution but which can be upgraded independently.
Of course nobody even knew so the default response is to hit that bottom arrow :D
Your entire reading of that is just wrong.
There is some discussion about the rationale at https://peps.python.org/pep-0594/#rationale
They won't remove them… but they regret having made them.
I think without, people would have just not used python.
To be fair, most things are missing from the official documentation. When I learned kotlin, I read through their official docs, and knew about most language features in a day. When I learned python, I constantly got surprised by things I hadn't seen come up in the docs. For instance decorators was (still is?) not mentioned at all in the official tutorial.
https://docs.python.org/3/glossary.html#term-decorator
https://docs.python.org/3/reference/compound_stmts.html#func...
But then I don't know how you're supposed to learn the features that are not in the tutorial. You can have a look at the table of contents of the standard library documentation for modules that might interest you, but that doesn't cover language features. Those are documented in The Python Language Reference, but that document is not really suited for learning from.
There are lots of websites and Youtube channels and so on, but you have to find them, and filter out the not-so-good ones which is not easy, especially for a beginner. I think there is room for some kind of official advanced tutorial to cover the gap.
There now -- https://docs.python.org/3/library/re.html#match-objects
re.Scanner looks more succinct though...!
And they're now just regurgitated things from the web, I've had novel ones generated fine (obviously you need to test them carefully still)
I've played with Antlr but never really vibed with it to be honest (and it was years ago)
use regex_automata::{
meta::Regex,
util::iter::Searcher,
Anchored, Input,
};
#[derive(Clone, Copy, Debug)]
enum Token {
Integer,
Identifier,
Punctuation,
}
fn main() {
let re = Regex::new_many(&[
r"[0-9]+",
r"[a-z_]+",
r"[,.]+",
r"\s+",
]).unwrap();
let hay = "45 pigeons, 23 cows, 11 spiders";
let input = Input::new(hay).anchored(Anchored::Yes);
let mut it = Searcher::new(input).into_matches_iter(|input| {
Ok(re.search(input))
}).infallible();
for m in &mut it {
let token = match m.pattern().as_usize() {
0 => Token::Integer,
1 => Token::Identifier,
2 => Token::Punctuation,
3 => continue,
pid => unreachable!("unrecognized pattern ID: {:?}", pid),
};
println!("{:?}: {:?}", token, &hay[m.range()]);
}
let remainder = &hay[it.input().get_span()];
if !remainder.is_empty() {
println!("did not consume entire haystack");
}
}
And its output: $ cargo run
Compiling regex-scanner v0.1.0 (/home/andrew/tmp/scratch/regex-scanner)
Finished dev [unoptimized + debuginfo] target(s) in 0.31s
Running `target/debug/regex-scanner`
Integer: "45"
Identifier: "pigeons"
Punctuation: ","
Integer: "23"
Identifier: "cows"
Punctuation: ","
Integer: "11"
Identifier: "spiders"
A bit more verbose than the Python, but the library is exposing much lower level components. You have to do a little more stitching to get the `Scanner` behavior. But it does everything the Python does: a single scan (using finite automata and not backtracking like Python), skip certain token types and guarantees that the entirety of the haystack is consumed.I posted this because a lot of regex engines don't support this type of use case. Or don't support it well without having to give something up.
That said, I don't usually use regexes for this either. Instead, I just do things by hand.
So I'm probably not the right person to answer your question unfortunately. I just know that more than one person has asked for this style of use case to be supported in the regex crate. :-)
if __name__ == "__main__":
block allows you to do this for a _module_, i.e. a single *.py file. If you want to do this for a package, add a `__main__.py`You can also use either of these throughout your code so that you can have
python -m foo
python -m foo.bar
python -m foo.bar.baz
each doing different, but hopefully somewhat related, things.This prevents a ton of import problems, albeit for the price of more verbose typing, especially since you don't have completion on dotted path.
It is my favorite way of running my projects.
Unfortunalty it means you can't use "-m pdb", and that's a big loss.
I'm relatively new to Python (used it for ~1 year in 2007/2008, again briefly in 2014 -- which is when I believe I picked this module trick up -- and then didn't touch it again until March of this year). It's made an impression on my team and we're all having a good time developing code this way. I do wonder, though, what other shortcomings might exist with this approach.
Or `breakpoint()` after 3.7 (better because the user can override pdb with ipdb or other with `PYTHONBREAKPOINT`)?
Just
python -m pdb -m module
EDIT: it support other options as well, like -c. That deserves an alias:
debug_module() {
if python -c "import ipdb" &>/dev/null; then
python -m ipdb -c c -m "$@"
else
python -m pdb -c c -m "$@"
fi
}python -m http.server is the most I have done.
If you have a fancy IDE feature, open a new python file, type "import pdb", use go to definition on pdb to jump to that file in the standard library, and read its main function - it handles -m explicitly :)
Multiline statements are not accepted, nor things like if/for
Even list comprehensions and lambda expressions have trouble loading local variables defined via the REPL
Are there workarounds? It would reduce the need for using IDEs. People who have experience with Julia and Matlab are very used to a trial and error programming style in a console and bare python does not address this need
Since I still enjoy a cmdline debugger more than a graphical one, I use ipdb, which doesn't suffer from the multiline limitation.
However, scoping issues with lambda and comprehension are actually a Python problem, not a pdb problem.
The scope problems are more fundamental:
The pdb REPL calls[1] the exec builtin as `exec(code, frame.f_globals, frame.f_locals)`, which https://docs.python.org/3/library/functions.html#exec documents as:
"If exec gets two separate objects as globals and locals, the code will be executed as if it were embedded in a class definition."
And https://docs.python.org/3/reference/executionmodel.html#reso... documents that:
"The scope of names defined in a class block is limited to the class block; it does not extend to the code blocks of methods - this includes comprehensions and generator expressions since they are implemented using a function scope."
This is a fundamental limitation of `exec`. You can workaround it by only passing a single namespace dictionary to exec instead of passing separate globals and locals, which is what pdb's interact command does[2], but then it's ambiguous how to map changes to that dictionary back to the separate globals and locals dictionaries (pdb's interact command just discards any changes you make to the namespace). This too could be solved, but requires either brittle ast parsing or probably a PEP to add new functionality to exec. I'll file a bug against Python soon.
[1]: https://github.com/python/cpython/blob/25a64fd28aaaaf2d21fae...
[2]: https://github.com/python/cpython/blob/25a64fd28aaaaf2d21fae...
It's "interact" though, not "ive".
You can exit it to carry on with the regular pdb.
Not just relative imports but also (properly formed) absolute imports. For example, if you have a directory my_pkg/ with files mod1.py and mod2.py, then
# In my_pkg/mod2.py
import my_pkg.mod1
will work if you run `python -m my_pkg.mod2` but will fail if you run `python my_pkg/mod2.py`However, the script syntax does work properly with absolute paths if if you set the enviornment variable `PYTHONPATH=.` (I don't know about relative paths - I don't use those). That would presumably allow pdb to work (but, shame on me, I've never tried it).
[0]: https://github.com/python/cpython/blob/3fb7c608e5764559a718c...
[1]: https://docs.python.org/3.12/library/sqlite3.html#command-li...
Not sure what you mean -- sqlite3 is the SQLite CLI.
The SQLite project provides a simple command-line program named sqlite3 (or sqlite3.exe on Windows)
that allows the user to manually enter and execute SQL statements against an SQLite database
or against a ZIP archive.I don't see why they would introduce _new_ code without annotating it, when that's clearly the trend for 3rd party libraries. From a quick look, it doesn't seem like it would be difficult to type either.
Though, I agree that typing everything is good. Especially when combined with a good typechecker like pyright.
grep -v 'test/' | grep -v 'tests/' | grep -v idlelib | grep -v turtledemo
Becomes: grep -ve 'test/' -e 'tests/' -e idlelib -e turtledemozipfile
Decompress a zip file:
python -m zipfile -e archive.zip /path/to/extract/to
Compress a directory into a zip file: python -m zipfile -c new_archive.zip /path/to/directorywww='python -m http.server'
Not near a PC now.
For example if your Ubuntu Linux desktop machine has host name foobar, and you run a http server on for example port 8000 then you can use your iPhone and other mDNS / Bonjour capable things to open
And likewise say you have a macOS laptop with hostname “amogus” and for example http server is listening on port 8081, you can navigate on other mDNS / Bonjour capable devices and machines to
On plasma it also installs a right click shortcut to share directory from dolphin.
I suspect it's related to relative path resource but never figured it out.
It may be a fantastic, well loved language that's exploding in popularity and the source of endless very high quality CLI tools... but. The absolute cheek!
We must only mention Python and Bash forevermore.
One of the biggest turn offs about Rust is the community.
This is even more fun on MacOS if you combine it with the pbpaste/pbcopy utils:
alias json_pretty="pbpaste | python -m json.tool | pbcopy"
That command will pretty-print any JSON in your clipboard, and write it back to the clipboard, so you can paste it somewhere else formatted! alias json-format="python -m json.tool | pygmentize -l javascript"Using modules (even if they are in the standard library) on the command line lets malicious code in the current dir take over your machine:
That is a feature, though, isn't it?
Edit: I also recall issues with wheels where .so's in unexpected locations would take preference over the .so files shipped by the wheel. I believe most of that should be fixed nowadays with auditwheel and hardcoded rpaths.
had to look myself. apparantly, it's like setting PYTHONSAFEPATH which prevents 'unsafe' paths from getting added to sys.path. new in 3.11
As opposed to line buffered, I assume? That sounds annoying, but why is it a security problem?
> no randomized hashes
I'm not up to date, but I think last I looked, I had the impression that randomized hashes didn't seem like they would fundamentally prevent collision attacks, just require more sophistication. Is that not the case?
> Seth pointed out this is useful if you are on Windows and don't have the gzip utility installed.
Okay, so instead of installing gzip (or just using the decompressors that aren't the official gzip utility but that do support the format and already ship with Windows by default[1]), you install Python...?
Even if the premise weren't muddy from the start, there is a language runtime and ubiquitous cross-platform API already available on all major desktops that has a really good, industrial strength sandbox/containerization strategy that more than adequately mitigates the security issues raised here so you can e.g. rot13 without fear all day to your delight: the browser.
The topic of this thread is the safety and security of running Python on (or merely next to) arbitrary content.
Even ignoring that: is "most people" greater than, less than, or the same amount as "all"?
So does the pendantry.
Seems to be a quite rare vector for exploitation.
Sure, on a multiuser system I might trick some other user into running such a command in /tmp and prepare that directory accordingly, but other vectors seem more esoteric.
https://www.gnod.com/search/?engines=af&nw=1&q=python%20-m
Even in Google's own repos. Starting any of those (no matter where they are stored) in a hostile repo would let the code in the repo take over the machine.
python3 -m http.server 8080
and then access the server on the receiving device through its IP address. It’ll show a basic directory listing, and you can download from there.One of these days it would be nice to make an unofficial Python reference book which documents these tools, hidden features (like re.Scanner!), and other corners of the stdlib or language.
An extra tool, but you don't really have to learn it, and if you're "always dumping JSON" and also use the command line a lot, you probably want to have it around anyway.
If I said 'Windows comes with the Edge browser' would you say 'I use two at work, one does but the other only has Internet Explorer'? Surely it's generally implied we're talking about things as they are, unless specified otherwise?
https://www.google.com/search?q=site%3Adocs.python.org+%22py...
I had the same experience with the tkinter ones - I thought they might be like zenity, a way to build simple UI elements from the command line. But they mostly just show simple non-configurable test widgets. The colour chooser could be helpful though.
So random and math are somewhat usable that way
$ python -m fire random uniform 0 1
0.5502786602920726
$ python -m fire math radians 180
3.141592653589793
$ python -m fire math e
2.718281828459045
Just running $ python -m fire random
will give you a nice "manpage" for your module as well. python3 -m http.server
I use it all the timeThat's why Python requires so many blog posts.