Calla – spatial video conferencing software based on Jitsi Meet
github.com
github.com
I used it briefly in 2009 and one of the things we tried to do with it was exactly this — spatial voicechat in a virtual environment with contrived "physical" modeling of spaces to adjust how sound travels. (This also reminds me of TeamSpeak's spatial integrations from around the same era.)
I know it's the pandemic that's inspired a ton of these projects, but they've been around in some form since live voice-over-IP came about and haven't really taken off in non-game applications.
What keeps blocking them from mainstreaming? I suspect maybe it's adding contrived physical interactions to the already high relative overhead of meeting virtually, but I'm not sure if that's just a problem for me. I could see myself enjoying something like this but not many others on my team seeing enough value to take the efficiency hit of having either zero or multiple open voice channels that do away with the abstraction.
[1] https://www.engadget.com/2008-10-22-massively-interviews-rap...
> What keeps blocking them from mainstreaming?
Most likely it is business model. Getting the spatial model correct is a lot of work, and making it generic enough to be reusable is even harder. (Both Second Life and Metaplace flirted with idea of being platforms that other services could build on.) It's easier to get some financing for it as a "game" (it's also useful to note a bunch of Metaplace ideas came out of experiments in failed MMO Star Wars Galaxy), but it's harder to sell it as a general chat application (much less platform) or especially as an "Enterprise tool" if it looks too much like a toy or a game. On the flipside it's harder to build it as an Enterprise tool first, because the impulse to keep things sparse/spartan/"profesional" also leaves a air of "lifelessness" with not enough things to do outside of/between meetings/during downtime inside of meetings, not enough of a feeling that it is a space to "inhabit" rather than visit briefly because a meeting required it.
These business model bootstrap hurdles are a criticism I often bring to a lot of the conversations about why the Cyberpunk "ideal" of a single shared VR space (including the one most recently popularized by Ready Player One) seems so extremely unlikely in the real world.
(Aside: Ready Player One especially suffers from a lot of basic game economy mistakes and likely shouldn't have survived very long in its world for very long, much less have taken over nearly everything including education and enterprise so much as it did. My reading of the book was that the game economy directly caused a lot of the collapse and dystopian economy outside of the game, which made the book better than the author seemed to have intended as further reading and other evidence suggests the author was not aware of how terrible the economics of the game were, they were just fun ideas thrown out for verisimilitude.)
I was thinking that the old 2d / move around 3-d kind of space chat from waaay back - worldsChat - you had an avatar and moved around a space station chatting.. I always wondered what held that back from being more popular / more used by more people..
Which brings me to the time of having all sorts of friends and family pinging on yahoo messenger for a while - then that was killed off.
Perhaps if worldsChat had this spatial talking kind of option it would of gotten more traction. I'm sure moderating and other factors also apply.
I think it's beautiful how it's "just" a wrapper on top of Jitsi Meet as it demonstrates how Free Software can (also) be used to experiment with and adapt user interfaces to specific and/or special needs instead of locking you into ever-more-bloated proprietary apps <3
A Zoom breakout room is nothing like that experience. This, or something like this, could very well be what is needed.
Can't wait to try this with my co-workers. One of the big barriers for remote casual conversation is the spatiality, since all video conference systems are focused on a single speaker.
Calla is a type of lilly. I name all of my projects after plants. Originally, I named this project "Lozya", which is the Hungarian word for "vines", because Jitsi is the Hungarian word for "wires". But nobody could pronounce it correctly, so I renamed it.
Incidentally, I work for a foreign language instruction company. I brought it up with one of our teachers and she didn't think it was an issue. She said that, in context, it wouldn't be read as "shut up", that it's kind of like "lead" (to guide something) and "lead" (a metal) in English.
Edit: And is it possible to self host?
No self-hosting planned, but private rooms with moderation controls are on the roadmap.
Products like these (and others like https://www.macro.io/) could exist as meeting apps, allowing people to to tailor rooms specific scenarios like one on ones, retrospectives, happy hours, etc.
This is more oriented toward professional meeting, but is part of the trend of doing more that just plain video conference. code https://github.com/retrolution/open-fishbowl and version on https://retrolution.github.io/open-fishbowl/?tableViewEnable...
I started this project as an experiment, half to see if it would work for my saturday morning tech meetup, half to see if I could make spatialized audio conferencing work in the browser at all, as at the time I was considering switching my VR app at work from Unity to WebXR (which, ultimately, I did).
The repository has been a little neglected in the last few weeks, but that's because I've been working on redesigning a few parts to make it work better in both 3D and 2D and haven't settled enough yet to commit the work.
I think the video stream in teleconferencing has absolutely zero redeeming qualities. People think it's there to be able to convey facial expressions and facilitate non-verbal communication, but I think it's a complete failure at that task.
For one thing, few listeners are actually looking at the speaker as the speaker is talking. They're most likely looking at themselves or some other thing going on with their computer, so their facial reactions are not really based on what the speaker is saying.
On the flip side, the speaker never gets to see people looking at them. Almost nobody looks at their camera instead of something on the screen, so the speaker never gets that "eye contact" feeling. Best-case scenario, you get a group of people trained to move the speaker video feed directly under the camera lens and they are diligent about making sure they are looking at the speaker. Even then, there is still a "20 yard stare" look to everyone. It also causes exhaustion as it puts you into a feeling that you're in an interrogation of some kind.
Additionally, it's such a narrow field of view for the camera. Non-verbal communication is more than just facial expressions, it's also body posture and standing distance. There are facial ticks that are also lost in the low-quality of the webcam feed, and the non-uniformity of every user's personal lighting settings creates an unnatural scenario where every person is lit differently than you'd expect, or from each other.
And finally, while teleconferencing has a lot of trouble with latency between when a person speaks and when the other people hear them, there is also a lot of latency between when you hear a person speak and when you see them. The audio and video feeds are not synced correctly.
By completely eliminating the video feed, conversations actually work a lot better. I get so many people who demand to have that video feed for the reasons they've been indoctrinated on, with little to no effort to even try audio-only conferencing.
And frankly, as a listener, I don't want the speaker to see my reflex reactions. I don't want them to see into my room. I really only want them to see the what I choose to let them see.
Thus, the avatars and the emoji reactions.
Would be nicer to just be able to point them to a webpage, but gotta take what you can get :)
I tried gather.town the other day for something similar but couldn't get it to work on Linux, couldn't detect my mic and camera correctly.
I am able to hear some background noise on the Calla room as I approach some others but I do not hear any talking and I do not know if others can hear me. I also can't get my webcam working, just a black screen and my webcam doesn't activate. It activated the first time I joined but now seems to have chosen a different camera.
I think camera/mic selection options need to be added for systems with multiple devices.
Really hope this gets refined!
sort of like ... can you have a map that blocks someone without letting them know you just distanced yourself? :)
Technically, Calla is just the library for driving Jitsi and adding spatialization. You can do whatever graphics you want with it. The graphical elements and the interactions are all up to you to build, separately from the teleconferencing.
[1] https://en.wikipedia.org/wiki/Croquet_Project
and later
[2] https://en.wikipedia.org/wiki/Open_Cobalt
?
One thing I've noticed with a lot of the "competitor"[0] apps out there is that they're more focused on the graphics than the audio. They're more focused on making a 90's-era RPG than on audio conferencing. I'm the opposite. I'm more focused on conferencing than the game. I get a lot of requests in the github repo for functionality in the game. The "game" is beside the point. It's just an exercise of the audio conferencing.
[0] I don't really see myself as in competition with gather.town or High Fidelity or any of the other apps because I'm not building Calla to be a startup. I love my day job and I won't be leaving it. Calla is just a component of a much larger thing I'm building.
That server cost seems very high, the same offering (2 vcpu, 4GB RAM) is €5.58/month at Hetzner for example.
https://github.com/capnmidnight/Calla/blob/master/Calla/doc/...
As for the server costs, I'm not going to use anything other than Azure. Most of the cost is bandwidth, but it's worth more to me to keep everything in one place and not have to worry about learning another PaaS system. From the other offerings I've seen, Azure is usually only more expensive at the most lowest of tiers. Once you start scaling up, the other vendors get just as expensive. So I'm fine with shelling out a few more bucks here and there to avoid having to use something else.
The video is 320x240 and one stream is about 600 kbit/s and audio is another 30 kbit/s.
So yeah, that works fine for a few users that have maybe 5 mbit/s upstream.
I built something similar as OP and chose a mesh solution for this reason even though it's inferior.
It appears that only big video chat companies like Zoom or Skype can afford to have a generous free tier, subsidized by their business offerings.
No one doing serious amounts of bandwidth pay outrageous big cloud prices. See: netflix.
What you got is a 200Mbps connection to your provider (and still that's probably a lie, it's probably shared before reaching their endpoint), afterward it's fully shared with every other customer... that's just how the internet is made, you can't have a dedicated 1 gbps to every single server, that just doesn't make sense.
Thing is, the higher the requirements, the more expensive it is to support, that's simple math... If you got 10 000 clients that download 1 gbps, you need 10 tbps, it's even worse if they are all on the same service, that connection won't support this, believe me.
1: The Telephonoscope was already imagined in 1878: https://upload.wikimedia.org/wikipedia/commons/8/8e/Telephon...
2: William Gibsons is often credit with the idea of "cyberspace" as in a virtual environment where people choose their own avatar and move around interacting with each other, with systems or data. https://en.wikipedia.org/wiki/Burning_Chrome#Reception
3: Gibson might have played MUD in the 70s: https://en.wikipedia.org/wiki/MUD1
I can only imagine that the charismatic or attractive person so going to get bothered/hit on/harassed 10x more in a virtual space where there's effectively unlimited space to be "near" to them and zero effort to do so.
Slightly off topic, is there any video conferencing software that doesn't use 100% of the CPU?
I hope Google/Microsoft/Zoom pay attention to this as well. And also Logitech, Bose, etc, so that we see headphones with accelerometers and gyros built in.
My first experience with spatialized audio in a chat system was AltspaceVR, when I first got into working in the VR space, when they interviewed me for a job (sadly, they didn't understand their own product enough and insisted that I'd have to move coasts to work with them). That was easily 6 years ago.
A couple of years later, I built my own VR-based spatialized chat system. It was way too early. WebVR was way too much friction for users[0].
Incidentally, Calla has all the necessary parts to be usable in a VR system. And I'm using it for that in my day job (VR environments for roleplay scenarios in learning foreign languages).
I don't remember why exactly I open sourced this project. It just seemed like the thing to do. I guess I hoped that other developers would find it useful and contribute bug fixes. That largely hasn't happened, so I'm mostly just focused on building the day job project for now. Once I get the next milestone done, I'll consider working on more features for Calla, but for now I need to pay the bills.
[0] WebXR today still kind of is, but it's a lot better than it used to be and I think some clever design in UX with something akin to 2-factor auth can help get around it.
I do really wanted to love Jitsi, but I've had frequent bad experiences. Every attempt I've had at a meeting has at least one person who's audio and video is bad to the extent that the rest of us can't really communicate with them.
I really do want to use it, but it always ends up with "Screw this, let's use Zoom"
'Use zoom and get screwed.'
https://github.com/jitsi/jitsi-meet/issues/4758#issuecomment...
It's not my fault that Mozilla is a shitty company that makes a shitty browser and that every other company that used to make a browser is now skinning Chromium. I just live in this world, I didn't make it.
You want better Firefox support? Put up or shut up. It's open source.