"River" detection in text
dsp.stackexchange.com
dsp.stackexchange.com
Ok, solving this kind of thing is rare, but still...
If you can't really get such stuff easily at the very first time does it mean you are a bad programmer?
Not quite.
The way I look at it, I would consider somebody good if he can read the documentation/theory etc and then come up solution to problems(read: can write programs). That way you know the guy can actually get some work done. It requires practice and time to get used to and comfortable to a new domain. That's natural friction you have to wear out. Regardless of whether its math, music, literature of whatever. Paying too much importance to factual stuff isn't of much use. What's more important is, can the person work his way out of the problem.
I checked out your bio, I understand you do math for a living and are obviously a little defensive when some one underplays the importance of your area of expertise. But the fact of the matter is though math is relevant, the areas in which its relevant are largely either rare or are already solved and presented to most programmers as libraries and frameworks. And if its not, simple analysis, a little reading and experimenting is sufficient to solve the problem at hand.
Your problem is not with programmers but with a term called abstraction. And fighting that is futile. The benefits vastly over weight any intellectual argument you can present against it.
EDIT: By the way after reading your bio, I have developed atmost respect for all the work you have done in your life.
Downvoted.
Your problem is not with programmers but with a term
called abstraction. And fighting that is futile. The
benefits vastly over weight any intellectual argument
you can present against it.
Utterly bizarre, to be claiming that a pure mathematician has a problem with abstraction.Side note, does the theory of juggling have a mathematical underpinning?
Don't care.
And yes, there is a great deal of mathematics underneath the structures of juggling patterns. Using them we predicted the existence of previously unknown juggling patterns, and have now used the math to create a proof that of all the juggling tricks of a certain type, we know them all (for some definition of "to know").
diagnose <-> disease
It clearly did not work as intended.
However, there appears to be such a depth of misunderstanding, or non-understanding, I feel compelled to try ...
Having read, re-read, and re-re-read your comment, let me say this. You appear to completely misunderstand the point. I'm not talking about mathematical ideas that have been taken, used, programmed, and made available through libraries. I'm talking about recognizing problems and techniques when they turn up unexpectedly, and in disguise.
Many, many times in my 30 years as a working programmer have I found elegant solutions to previously intractable problems, purely because I recognized some sort of weird mathematical structure lurking underneath. It's my specialty to go into a situation where the domain experts have worked to solve a problem, sometimes for years, to have them explain the problem to me, and then to bring different insights and tools to bear on the problem. I'm not necessarily more clever, or more capable, I just have a different toolbox. And when the problem at hand is unusual, my toolbox is particularly powerful.
You said:
Somebody outside devops would consider pervasive use of sed, awk
and other unix text processing utilities line noise. Yet such a
thing would come naturally to that person working on it daily.
I have no idea what point you're trying to make here. If you can't really get such stuff easily at the very first time
does it mean you are a bad programmer? Not quite.
Of course it doesn't. Different programmers have different skill sets, and different experience. Someone who doesn't know databases or web programming may be a wonderfully productive programmer in another area of specialization.I have no idea what relevance this has to my comment.
The way I look at it, I would consider somebody good if he can read
the documentation/theory etc and then come up solution to problems
(read: can write programs). That way you know the guy can actually
get some work done. It requires practice and time to get used to and
comfortable to a new domain. That's natural friction you have to wear
out. Regardless of whether its math, music, literature of whatever.
Paying too much importance to factual stuff isn't of much use. What's
more important is, can the person work his way out of the problem.
Perhaps. My point is that people often dismiss mathematics as useless to programmers because they can't see how this lemma or that theorem can ever be relevant. My point is that the methods of thought trained by the study of advanced mathematics has, in my experience and others', regularly led to deep insights in apparently unrelated problem domains. This ability is not something you can "just read up on." This isn't a case of reading documentation and/or existing theory - this is a case of seeing something as being a disguised example of something else. I've lost count of the number of times I've recognized a matching problem, or a topological separation, or an abstract vector space, or a group acting on a metric space, and that recognition has subsequently led to a new solution. I checked out your bio, I understand you do math for a living and are
obviously a little defensive when some one underplays the importance
of your area of expertise. But the fact of the matter is though math
is relevant, the areas in which its relevant are largely either rare
or are already solved and presented to most programmers as libraries
and frameworks. And if its not, simple analysis, a little reading and
experimenting is sufficient to solve the problem at hand.
Then, with respect, your experience appears to be limited, and your opinions parochial. Your problem is not with programmers but with a term called abstraction.
And fighting that is futile. The benefits vastly over weight any
intellectual argument you can present against it.
This is just bizarre. I'm a pure mathematician - how can you suggest my problem is with abstraction? How can you possibly think I'm offering any argument against abstraction?Yes, I agree my experience is limited. In fact now that you mention you have been a programmer for 30 years. I now realize I have much more to learn and understand, Its possible I have simply don't even know what you know and how deeply you understand things.
What advice would you give to some one like me.
> I meant the problem with 'your argument'.
To recap the discussion:==========
Chirono: I find this comment interesting given the recent ... talk about not using maths puzzles in programmer interviews. This is perhaps an interesting real-world case that demonstrates how a mathematical outlook leads to much cleaner code.
thomasz: The secret sauce here is not the ability to solve "maths puzzles", but specialized domain knowledge.
I replied: That's exactly the wrong way round. You can acquire domain knowledge, you can have a domain expert assigned to you to work with you. What you can't acquire, on demand, is the ability to think in the ways that mathematics gives you. That requires extensive training and practice, and recognizing when obscure bits of theory are applied is something that doesn't come overnight.
==========
So my argument is that while domain knowledge is a good thing, and (jacquesm eloquently argued) often necessary to solve a problem, it usually can be acquired as needed through reading, study, and discussion with local experts. The ability to recognize, let alone solve, problems that are actually obscure mathematics in disguise cannot be acquired on as "as needed" basis. This is why knowledge of advanced mathematics needs to be gained early, and if you don't have it by the time you are a practising, working programmer, you are unlikely ever to acquire it.
Not having acquired such skills does not make one a bad programmer. I know very, very few programmers who hold advanced degrees in mathematics, and yet the world is full of programmers who can "get the job done" efficiently and effectively. Equally, having skills in advanced mathematics does not, of itself, make one a better programmer (and certainly not a better person.) I know many PhDs in mathematics who simply cannot program effectively.
But it does make one a programmer with different skills, and those skills can enable one to solve problems others can't, and sometimes to produce solutions that are more elegant. Not always, not every problem requires such skills. But if it does, by the time you find the problem, it's too late to acquire them.
> What advice would you give to some one like me.
I don't know what you like, what skills you already have, or what you want to end up doing. I studied math because it was fun, exciting, and I was good at it. Computers at the time didn't exist in the form they do now - maybe I would have done computing. Impossible to say.So I can't give advice - there would be a real danger of simply trying to turn you into a clone of me, and that's a bad idea. I can tell you that understanding algorithms, time-complexity, calculus, vector spaces, and point-set topology have all been of direct and immediate use in my work.
And if you'd like a simple puzzle, here:
Suppose the demand for a product falls linearly as a function
of its price. Show that if you price the product to maximise
your profit, you will have less than half the possible market.
That turned up a few years ago in a discussion with a customer.The majority of development roles do not need any more mathematical knowledge than any other job.
These methods work for this particular image, and give an idea about the possible solution, but the thresholding and morphological operations are not scale invariant, and very sensitive to light changes and noise.
Any algorithm that can detect rivers in text from any image, acquired by some device, such as mobile phone or a scanner, in any light circumstances, will run hundreds of lines long.
Still, that's a pretty expensive fitness test, and I wonder if there's a more elegant and efficient approach than Monte Carlo methods.
If it was, say, a 4MHz machine, you both could be right.
It gets to 3.14 within a second, about 30 seconds to 3.1415.. and Ruby is surely doing this in 100x the time C or Java would.
Update: Also did it for JavaScript - https://gist.github.com/peterc/5019804
OK, so I was going to just do the unit circle thing. Then I got stuck in optimization mode. God I miss performance graphics programming.
It's a nice example but almost useless in practice.
Rivers stand out more in blurred text, even in central vision.
Some proofreaders would note them should they find / notice any but some wouldn't.
Now that I think of it, using OP's findings it would probably wouldn't really be too hard to write a "plugin" taking screenshots of Quark XPress / Adobe InDesign / LaTeX preview and showing "rivers" in a HUD on top of the screen-rendered pages. This would make chasing "rivers" quite easy.
Relevant xkcd: http://xkcd.com/1015/.
These sorts of image processing questions are incredibly entertaining imo (and I don't mean it in a bad way).
I simply don't know what a Hough Transform is and rather than Wikipedia, I would rather a coursera course on image processing - it's to get things in context that matters.
As I get older i am not afraid of hard work, just afraid of exploring in every direction - the waste of time and effort is the problem, time and effort finding out what to learn rather than the practise of learning
Oh, just answered my own question - coursera!
If TeX is already using dynamic programming in order to improve the visual appearance of the words, I imagine that the same can be done for the space between the words without having to resort to image processing of the TeX document rendered as PDF.
You could even do some sensible analysis per glyph.
width of base: A
vs width of top: V
IMHO the machine learning solution could be better, as rivers are defined more aesthetically than functionally.
Your next problem is that TeX doesn't know what glyphs look like, just how much space they take up. If the "riverbanks" are made up of As and Vs, the river will be extremely pronounced. But it may not be noticeable at all if they're made up of Ms and Xs. Worse, it will also depend on the font in play. But TeX doesn't know what the glyphs look like while it's adjusting the whitespace, so you now need to introduce a feedback loop between the eject-page phase and the paragraph layout phase. I'm under the impression you'd be jumping over three or four phases to do that, because TeX lays out lines first, then paragraphs, then emits pages.
This is a huge amount of work. How big a problem are rivers?
Oh the memories when I used to write (and typeset) books using LaTeX and Quark XPress: "rivers" had to be tracked down manually, by eye-ball searching. You'd basically "blur" your vision a bit and quickly scan through all the pages of the book. I didn't take that long and I don't write books anymore but I'd still be curious as to when that technique is going to be implemented (maybe in InDesign which I never used!?).
So nowadays, I don't manually check each page for rivers anymore. When designing the page layout, I decide on the 'color' (the density of the paragraph) that works best, so that rivers are avoided, but I leave it at that.
I understand how river detection can be a fun mathematical puzzle, but it's moot. Adobe built the Multi-line Composer into InDesign since version 1.0 (1999). And in my experience, with every version, it gets better.
From MacWorld's review of InDesign 1.0:
"InDesign's text features will appeal to designers and production people tired of the drudgery of manual copyfitting and kerning. The Multi-line Composer feature can calculate hyphenation and justification settings by examining an entire paragraph (or as many lines as you specify), instead of just a single line, to create better-looking text. In the process, it notably reduces the amount of manual tweaking necessary and will be especially helpful if you're trying to avoid multiple word breaks in a design with awkward text wraps. Similarly, the program's Optical Kerning feature does its best to find optimal character spacing, even if you've mixed different type sizes and faces. "
http://www.macworld.com/article/1014955/k2.html
http://blog.paragonpress.net/2010/08/04/adobe-indesigns-para...
That's false. Almost any printed publication uses justification and hyphenation by default: books, magazines, scientific papers. The big exception is websites.