Koders: Search Over 3 Million Lines of Open Source Code Around the Web
koders.com
koders.com
An example: I searched for "helper_method" (a Rails method), and found these (bracketed what was bolded as the search hits):
class.Scheduler.php:
// Holds [method]s for ...
Editor.py:
from Models import Editor[Helper], Controllers
AutShare.h:
// IShare [method]s
and then I get over a dozen literally-identical files called "helper_method_example.rb" (tests with these results: require File.dirname(__FILE__) + '/spec_[helper]'
describe "a context with [helper] a [method]" do
def [helper_method]
and the rest of the files contained this after those lines:
"received call"
end
it "should make that method available to specs" do
helper_method.should == "received call"
end
end
There are about three pages of results like that, or the same wrapped in a `module HelperMethodExample`, with identical contents otherwise.I will happily pay money to anybody who can build a code search service that is actually functional.
It doesn't search code repositories, but rather tutorials and question/answer sites like stackoverflow. Right now stackoverflow dominates the results but that should change as we index more sites.
I ask as I am doing this but have not considered a business model for it as yet.
I wonder if you could even build a market off open-sourcing the engine, thus allowing people to improve it for you, and run a paid service that searches GitHub, Google, etc for results. Free accounts for contributors that fix bugs or provide speed / feature improvements?
I dunno, just throwing ideas out while they're percolating, I haven't given them a whole lot of thought.
I had also considered the GitHub model. I will be making portions free-software as the whole thing is built off it anyway.
If nothing else I will be going free for quite a while as I believe that there should be a code search repository available online.
For the record it will live at http://searchco.de/ I was in the middle of pushing out my second release when Google made its announcement so its looking a little worn at the moment.
Why we are moving from .NET to Java technology is easily addressed: the .NET implementation was built by a company we acquired; we are a Java shop. We can invest in and expand the code search engine and site more effectively using technologies familiar to our developers.
We haven’t yet upgraded Koders to the latest software version. But we’ve done quite a lot of work beyond just a vanilla language port. For example, search results will be filterable to a single (or multiple) projects as mentioned in one of the comments and by language(s) as mentioned in another. We have noted the feedback regarding duplicate results, underscores, quoted searches, etc. and will put these in our backlog.
In preparation for the debut of the new Koders, we are currently doing scalability testing to ensure that performance of the re-launched site is up to par. We will index a lot more code than is currently there. Our test instance is now running with more than 7 billion LOCs in-house, about 2x the size of the current Koders.com index.
If you’d like to download and try the Java-based implementation that will be deployed to power the new Koders, it’s at http://www.blackducksoftware.com/java_implementation_downloa... .
We look forward to making this a great code search resource. Again, thank you for your interest!!
"this sucks" kinda works. :)
Hopefully they can fix that, just having a bunch of source code indexed, or even available to be indexed if it turns out the current indexing is no good, is a good first step. But it seems like it needs work right now.
Interesting. I'd love to hear more about the rationale behind this move. I can understand moving from .NET to a dynamic or functional programming language, but why move to what is effectively a peer?
And in either case, moving from .NET to Java would seem more rational than moving to a new functional language because they are so similar and wont require as much work. I wouldn't be surprised if there even are automatic converters.
There are always other "reasons", but they are usually just straw men. Users don't care what your site is written in.
1. Regex search support. 2. Browse full file. 3. Filter search to single project.
This is based on what I observed on the Reddit and HN discussions when it was announced. I can see from the below some people would like grouping of identical results which is fairly easy to add so no worries there.
Any other thoughts?
Regex search doesn't seem like a requirement here; for searching code, it usually suffices to include symbols when you index. Might also help to have some language-specific text classifiers, so you can say "search comments" or "search strings" or "search code".
Browsing the full file seems quite important to check the results. For that, you'll want line numbers, per-line anchors, and basic highlighting support for popular languages, in order of priority.
Filtering to a single project doesn't actually seem that helpful; if I know what project I want to search, I can clone it locally and grep it myself. Code search engines help with questions like "does anyone use this API?" or "how have other people processed the data from this library?", which need to search all projects. Filtering by language might help, though (to avoid problems with APIs provided in umpteen languages).
print_r
During the indexing process it would take print and r as separate words. Any search for "print_r" should break up the same way, and assuming the search is phrase heavy return any document with "print_r" near the top anyway.
I think code search is very useful and I've come near to trying to construct a searchable database just for personal use. It's an itch I've been tempted to scratch. But it would require some time.
So far I have mostly everything working, its just pulling down enough repo's to make it worthwhile.
I will soon have code search as per Google Code Search. I just need to finish pulling down several thousand more repos before pushing it live.
For the moment however look at the following,