Bible Semantic Search
share.streamlit.io
share.streamlit.io
A clear example is the query: "homosexuality"
This returns:
> James 2:3 - And ye have respect to him that weareth the gay clothing, and say unto him, Sit thou here in a good place; and say to the poor, Stand thou there, or sit here under my footstool
It's clearly seeing "gay" but is unaware that the meaning has changed.
A classic issue when applying a ML model out of domain.
https://wiki.crosswire.org/Alternate_Versification is a decent starting point for looking into the topic, with SWORD’s canon_.h files fairly tolerable for showing the number of verses in each chapter. Unfortunately, SWORD has never gone as far as doing proper mapping* between versifications for some reason—they have some basic mapping somewhere or other, but I can’t remember offhand where or what it is, as I only briefly looked into it five years or so ago.
The most significant differences occur when you switch languages (there are quite a few differences if you switch from English to French or to many Indic languages), but there may be some differences between translations within a language too, e.g. SWORD’s NRSV versification has an extra verse in 3 John and Revelation 12 compared to its KJV versification.
I used to distribute a BibleGateway-ripped copy of NIV as a plugin for the open source GnomeSword project and managed to upset both Harper Collins and the open source scholar community. It was absurd and I refused to stop. I dared them to sue me for sharing the Bible as open source... the headlines would write themselves.
FWIIW they never took the bait, but expect empty threat Cease and Desist letters from the Harper Collins legal team.
ESV or CSB, please...
How hard would it be to train it on the range of common modern translations? It would be an interesting stress test of the models to see how close searches in different translations are - they're theoretically all communicating the same thing, with different styles and emphasis, but I'd expect a lot of that to fall out in the semantic search (at least, if it were working properly).
You could grab a range of English translations ranging from "very literal" to "thought for thought" (The Message applies here, and I'm not even sure it's thought for thought), do various searches, and see what the overlap in results is.
In any case, very neat project... concept. :/ It appears to have fallen over, all I get is "Please wait..." when I try to access it. Even without my usual web filters interfering. I think.
I haven't personally read much of ESV or CSB, just curious, why the preference for those?
[1]: https://en.wikipedia.org/wiki/Dynamic_and_formal_equivalence
Edit: Actually, better yet: https://crosswire.org/sword/modules/ModDisp.jsp?modType=Bibl... has licensing info for each listed item
As someone else pointed out, the WEB translation is public domain so it might be a better option.
This is awesome. Thank you for building it.
One of the things that makes KJV version useful for scholars is that since it's been the "standard" for hundreds of years a great deal of other work references its structure. People use it because of this. It's less to do with it being a fabulously accurate translation or whatever. It's just newer bibles don't have this wealth of history and documentation that references it and much of them are pretty expensive to license, while KJV is public domain.
For example "Strong's Exhaustive Concordance". If you get a version of KJV with "Strong's Numbers" you can cross reference words and phrases in their original languages (greek/hewbrew/etc). This way students can understand some of the original meanings that go into difficult or disputed passages.
Also there is a large number of commentaries of all sorts of different types that reference specific passages.
Besides that KJV is just mostly valued in Protestant Christian dialects of Christianity. Other Christian religions such as various versions of Eastern Orthodox have different numbers of books or will arrange things in different orders. There have been different attempts to past scholars to arrange things in more chronological order, too.
This makes the Bible fairly unique when it comes to literature. Each verse of text can have dozens of different "back links" and "references".
so if somebody searches for the subject "Homosexuality" it will get hits in various commentaries. Those commentaries all directly reference verses in the Bible.
So you could show the version found and why it was selected. That way a reader would be shown "These authors think this verse is has to do with homosexuality" and they could click through and find out the justification for this, different translations, what those translations are likely based on, what other Christian sects feel this verse means, and so on and so forth.
I don't know if it would be useful for you, but there is a "Sword Project" that collects and cross references different Bibles and bible resources as well as tools.
…and it's gone.
urllib3.exceptions.ProtocolError: This app has encountered an error. The original error message is redacted to prevent data leaks. Full error details have been recorded in the logs (if you're on Streamlit Cloud, click on 'Manage app' in the lower right of your app).
Traceback:
File "/home/appuser/venv/lib/python3.8/site-packages/streamlit/scriptrunner/script_runner.py", line 475, in _run_script
exec(code, module.__dict__)
File "/app/bible-semantic-search/app.py", line 50, in <module>
query_results = controller.query(
File "/app/bible-semantic-search/controller.py", line 127, in query
results = self.index.query(query_emb, top_k, namespace)
File "/app/bible-semantic-search/pinecone_index.py", line 63, in query
return self.index.query(Seems kind of appropriate. Obviously it is easier for a camel to thread the eye of a needle than to get a semantic search on the Kingdom of Heaven.
Explanation when you click “Wat this?”:
> This is a Streamlit app I prototyped for performing semantic search on the King James Bible. It conducts full text search as well as semantic search, which is useful for surfacing passages that are similar in meaning to the query, even if the passages don't explicitly contain the query keyword(s). Suppose you wanted to bring up all verses that reference the infamous snake that tempted Eve. In a traditional keyword search system, searching for 'snake' wouldn't yield any results because the KJV uses the term 'serpent'. A semantic search system would take that 'snake' query and retrieve the relevant verses that contain 'serpent' as well as similar verses like ones about reptiles.
> Under the hood, I've generated vector embeddings of every verse in the Bible using SBERT (https://www.sbert.net/), and stored those embeddings in a vector database called Pinecone (https://www.pinecone.io). Every time you submit a query, it's converted to its vector representation using SBERT. That query vector is then sent to Pinecone, which performs an Approximate Nearest Neighbor (https://www.pinecone.io/learn/what-is-similarity-search/) search, retrieving the top n verses that are the most semantically similar to our query. The verses returned are ranked in order of most to least similar.
(Full disclosure: I work for Pinecone, but I have no connection to this demo.)
In terms of this project, I would not have chosen pgvector. I don't want to deal with the PITA that comes with manually setting up and self hosting a vector database. When I'm building a demo or a prototype, I care about speed of development, which lets me more effectively explore the possibility space of whatever I'm building. I'm not a database admin, so when I deploy the project, I don't want to do database admin tasks. The ease of use of Pinecone's API lets me move fast, and it was very intuitive to learn. One downside of Pinecone is that although they have a generous free tier, their next highest tier is $50/month (and that's the low end of that tier). This is unfortunate for solo devs like myself who are likely to graduate from the free tier and would be willing to pay a bit more for a higher tier, but find the $50/month plan to be overkill. I think they're trying to target startups with that $50/month plan.
(I also use it to explain type checking to friends, given the number of times I've looked up a Hebrew word in the Greek dictionary, or vice versa).
- Precision@2 = 50%
- Precision@5 = 20%
- Precision@10 = 33%
Relevant results:
2) Philippians 1:21 - For to me to live is Christ, and to die is gain.
6) Ecclesiastes 3:13 - And also that every man should eat and drink, and enjoy the good of all his labour, it is the gift of God.
10) Ecclesiastes 2:17 - Therefore I hated life; because the work that is wrought under the sun is grievous unto me: for all is vanity and vexation of spirit.
For example I expected Luke 9:13 as a result for the phrase "you feed them" or "you give them something to eat." The King James reads "give ye them to eat" but the app never found it.
[1] https://github.com/chrislee973/bible-semantic-search/blob/ab...
totally missed when they made a sharing platform.
https://www.biblegateway.com/versions/World-English-Bible-WE...
Since the author of this app is based in the United States, he should be safe… we have a bit of precedent for ignoring royal prerogatives here :)
NKJV is more readable for people who grew up reading blog posts instead of old books and generally retains the poetry of the KJV, but it's under copyright. I think it would be allowed for a use like this though, because they explicitly allow (as do most translations) quoting up to a certain number of verses.
EDIT: Maybe not, here's the rules: https://www.harpercollinschristian.com/sales-and-rights/perm...
Seems like you'd fall afoul for the % of total text requirement.
>
> Mark 1:25 - And Jesus rebuked him, saying, Hold thy peace, and come out of him.