How and why we teach non-engineers to use GitHub at Thread
thread.engineering
thread.engineering
Not only is our website easier for non-technical folks to update, but that means we have built-in transparency for much of the business. There are private repos for things that don't need to be aired company-wide, but public is the default.
I really like this.
When it comes to productivity in the enterprise I'm quite fascinated by the intersections of communication immediacy versus persistence. We like things that are immediate when we speak, but like things that are persistent when we're digging around years later.
A searchable per-theme repo with searchable issues and dialog sounds like a dream. I dunno if github has this built-in, but their integration API is very nice so having direct email integration and notifications or even local notification clients should be highly solvable.
If the non-engineers are just using git to upload and version json files full of data without modifying the code that reads them, uploading to s3 seems like a much safer and easier way to accomplish the same goal (without building CRUD interfaces to handle it).
I can't see how any of that is worth the extra time spent teaching git, or the extra complexity that comes along with merging your code and data into one repository.
Git, and GitHub, at the level we teach them, are not very complex.
However, if you see github.com as a proprietary UI for git, it looks almost exactly like putting your files on S3 with some CRUD UI around it.
Not to mention that mixing data and code into the same versioning system adds complexity that I don't think is maintainable.
There is a reason we tend to version data separately from code.
As always, it depends on the data. Sometimes git is fine. Sometimes you need S3. Sometimes Postgres. There is no one size fits all solution.
Polluting the commit logs with thousands of extra commit messages if you're dealing with fast changing data for one.
If a developer is working on a "data file" that happens to contain logic (its going to happen no matter how diligent you are about keeping them separate), dealing with unnecessary merging from the constant edits is another.
>That is perfectly fine to store in the tree and I think the author was talking about that sort of data.
The individual datasets may be small, but from reading the authors comments here, the scope of what they're storing this way seems to be way beyond what I'd feel comfortable managing this way.
We’ve used this approach before we were able to intergrate required CMS, and it was fine. Our Product Owner would go into the YAML files in Github, click "Edit", make the changes, then hit the big green ”Create Pull Request" button. GitHub actually makes it fairly easy to do this flow and comes with sane default commit messages.
The article explicitly mentions that they teach them these things.
No it isn't. This is entirely a sales problem.
You may not have noticed, but engineers typically can't use distributed branching either. The volume of tutorial and "i'm in trouble, plz hlp" material out there for Git (and really every other VCS) is substantial. Everyone has a story about how a certain workflow got them into trouble on a project.
The thing is, collectively we see a huge value in it, so we make the learning investment to use it. It's only after we learn it (and the pitfalls to avoid) that it's easy.
Novices don't use it because they haven't had the value proposition made to them yet. If you can reframe the argument for why, you'll get them using it.
Sounds like it's not just a sales problem to get novices to use distributed branch management correctly.. but maybe a design problem with the current tooling. If it requires significant experience to use a tool correctly, then it is impossible for novices to use as you must turn them into experienced users.
My original comment was just to point out that GitHub is another UI on git repos, and thus might work around some of these design flaws.
Has anyone thought about doing this for content/marketing/design people? The tools that they use are hard too, but they know how because they were sufficiently motivated to learn them.
It's a big mistake for engineers to think that their jobs and the tools they use are significantly harder than in other fields.
I completely agree with everything you state in this comment, especially your last sentence. I have a strong sense that we have similar opinions but different ways of stating them.
Hmmm... ok, I'd agree it's a mistake to simply assume this. But I'm not going to dismiss the possibility, either. Consider this thesis: there is a constant degree of churn in software development that adds significant complexity to the task of software development relative to many other fields.
Is keeping up with javascript framework churn par for the course in other fields? It's very difficult to make that call, since few of us have worked in a wide variety of fields in a meaningful way (I worked in a law firm right after college, but I wouldn't say I can meaningfully comment on the practice of law). But I have a sneaking suspicion that our jobs actually are pretty difficult. We deal with what I feel comfortable calling a hellacious degree of technical complexity at times.
Since then, the "noSQL" movement occurred, and while I'm delighted to see a resurgence of interest in SQL, mongo, Cassandra, various graph databases, and other non-sql data stores do have valuable uses (I think some people just went too far against SQL and the relational model, which also remains exceptionally valuable). I didn't name all of them, but while there was value in there, some of it was churn and adopting technologies for the hell of it. But even when it wasn't churn and was legit, that's still a lot to learn.
Also, spark, hive, Hadoop, and a lot of storage and infrastructure has moved to a cloud environment (AWS is the biggie, plenty of others). Again, not necessarily bad tech, but a lot to learn.
Also, in my own particular field (this may not apply to all backend developers), machine learning did become a big factor in work, especially prediction systems. I did a lot of math in undergrad and grad school, but there was still plenty to learn there, including dusting off the old math books and re-learning how gradient descent works for neural nets, logistic regression, and then other approaches like bayesian methods, random forests, and so forth.
Also, in this time, the middle tier technology you needed to know also changed - to a large extent for the better, but C++ gave way to Java, which grew in to a mountain of a mess with EJB and other crap, then ruby and python, rails and Django (for the record, I loved ruby and rails but don't work in it much anymore). If you want to do distributed computing, you can use python, but you really might want to consider learning Scala.
Now, obviously you don't have to learn all this, you can specialize, and I wouldn't call it all churn (then again, I wouldn't call all the javascript framework stuff churn, either). But to work on the backend, you probably did have to keep up with a lot of that.
And like I said, I couldn't tell you what it's like to keep up with other fields... but yeah, I'm pretty comfortable saying that keeping up is pretty hard on the backend as well, and that learning these tools is no easy thing.
The tool isn't the skill, and honestly I don't want to work somewhere that values "x years using y" over "thoroughly understands concepts behind/around y".
I've been in IT since the late 90s, only development the last 5 years, but knowing how protocols work and what makes systems/architectures reliable gets way more play in my day to day job than knowing whatever the current tech is.
Git really ought to be treated as a low level tool that is called by a higher level tool that assembles content management workflows designed by developers that can be used by developers and non-developers alike.
Non-developers ought to be able to make 'commits' with git and pull requests without even realizing that git is the underlying tool.
Developers should just script their most common workflows and then use them 99% of the time.
Unfortunately the tooling that sits atop git is largely all shit and people have gotten it into their heads that one needs to have a deep understanding of git's quirks to be considered a 'serious' developer.
> Non-developers ought to be able to make 'commits' with git and pull requests without even realizing that git is the underlying tool.
Absolutely. By why 'ought'? What's stopping you?
Git has an API, lots of tools use it. There are commercial products that wrap git for designers and such. Many of my systems are GIT and Github enabled (Octokit), and the users are none the wiser.
> Developers should just script their most common workflows and then use them 99% of the time.
I am almost positive this already exists in many various forms online (on github in particular), for personal needs. I know a significant number of projects script checkins, builds, etc.
It breaks down, IMO, because of environmental constraints. You're trying to maintain alias sets across machines, on colleagues machines, or handle tricky merge issues per-project. At the end of the day the APIs breadth is required to handle corner cases, and familiarity with the breadth makes a lot of noise 'straightforward'.
I don't know of any decent tools. Something that would give an end user a nice UI to let them edit a markdown or JSON/YAML file and issue a pull request following a team defined workflow, for example? Or a WYSIWYG markdown editor that can issue pull requests following a team defined workflow at the click of a button?
>Git has an API
I noticed.
>I am almost positive this already exists in many various forms online
Most teams I've worked with don't have anything. One of the first things I do when I join a team is figure out their workflows (typically baked into people's heads) and get them down in a script.
>It breaks down, IMO, because of environmental constraints. You're trying to maintain alias sets across machines
A) git aliases are a poor substitute for a decent workflow program.
B) A lot of workflows integrate a lot of different things - e.g. by convention, branch names on new tickets often contain the JIRA ticket number followed by a slugified name. That means to properly script a workflow you need to make an API call to get a ticket number and slugify the name and use it to create the new branch.
I'm not against automating it, but introducing magic seems like the wrong approach to the problem.
(I mean for developers)
FWIW, that is how git is organized: low-level "plumbing" commands for the content-addressable file system, and the high-level "porcelin" commands for a VCS. Moreover, most git commands have human and machine readable variants, to facilitate abstraction.
no, people are usually lazy about being extra precise about tracking transaction state of any sort (in conversation, finance, rational discussion, etc), and moving this imprecision from a linear sequence into a tree and from there into a forest only compounds things.
most of the problem with (d)VCS is that people are used to fuzzy transactional thinking, and furthermore, need to learn a new vocabulary to interact with it.
Letting people who are editing templates use git seems perfectly fine to me. Using git to version constantly changing data is what I think isn't maintainable.
I think at some point that polluting the commit logs with data changes, dealing with merging issues, and training everyone who wants to change data on git imposes a hard limit on scaleability for this approach.
You may never hit that limit, but I don't think this would work for very large teams, and I suspect that if you git to the point that it stops working, it will be very difficult to sort it out.
As for the commit pollution, I don't see this as much of a problem. In the long term we're always going to have lots of commits!
For instance, I'm thinking of products like Conveyor: https://wildbit.com/blog/2017/02/14/our-next-product-conveyo...
Forestry has been a big help to us with support and quick turnaround for feature requests and it's been pretty smooth-sailing for our content team.
For internal documentation/knowledge-sharing, we started a MediaWiki.
Unfortunately neither of these work for obscure internal data models that we have, or the time to integrate one isn’t worth it for an A/B test we might throw away next week (although if we decide to ship the A/B test we would take that time to integrate a CMS or add an editing interface on top of our DB).
Throwing them away is a tremendous waste of knowledge. We found this to be true with our marketing and that was the reason for the push to setting up a Wiki. Even if we don't retain code around an experiment, we have a documentation format for recording what we made, what our expectations were and our results.
For example, we want to test showing a set of images on a page. Some users get set A, some get set B, some get none, and images need to be changed weekly.
The list of image URLs might be in a file in git, and get changed by a product manager each week for the duration of the test.
The A/B test would be run by our A/B testing service, which stores historical tests and reports indefinitely.
We’d then come and move the URLs into the database and add an editing tool later on if we decided to ship the A/B test.
We use the solution in the article for things like the current delivery estimate, which changes twice a year and is a Python timedelta object (example, but more complex than that) or adding new tags to our support ticket system (not user facing). For those, we don’t want or need to spend the overhead of making an editing interface.
As for integrating another CMS, the whole point is that it’s just Python files in the repo, zero effort for developers, very little effort for everyone else, when editing is infrequent.
That said, I don't question the benefits of version controlling data, and the usefulness of GitHub's web interface for quick and dirty commit work.
I am surprised that a quick search doesn't seem to turn up someone already trying to build a NetlifyCMS-like tool for inside the Django Admin. That seems like a natural fit for Django's CRUD philosophies.
This is okay now, but could become a major pain when your company grows and Nancy in Accounting who is an aspiring writer starts making _hundreds_ of copy changes with pull requests every week. You're going to really wonder where all your engineering time went when suddenly you have a dozen Nancy in Accountings...or a hundred.
If you do that, you are communicating to your employees that everyone is welcome to make contributions to the product. If not, why the f* would you do it?
I have access to pretty much every Confluence page at work, including ones other people set up for executives to look at, but I don't go around making changes and editing those pages just because I have access.
People aren't robots. You can hold them accountable for their actions and they can make their own (good or bad) choices.
https://github.com/crdoconnor/strictyaml
I wrote it partly to make it easier for non-programmers to maintain structured data - which seems to be your use case. YAML is cleaner and easier to maintain than python dicts/lists/etc, and with a strictly validated schema you can trigger failures caused by accidental insertion of invalid data at an earlier stage.
Python has worked fairly well for us so far though, we have tests that would detect invalid syntax, and importing a Python file is easier than parsing and doing the error handling on YAML/JSON.
e.g.
import strictyaml as sy
MY_TRANSLATIONS_DICT = sy.load(
Path("translations.yml").text(),
sy.MapPattern(sy.Str(), sy.Map({"French": sy.Str(), "German": sy.Str()})),
).data
translations.yml:
Hello:
French: Bonjour
German: Guten Tag
Goodbye:
French: Au revoir
German: Auf wiedersehen
It would negate the need for most of those tests checking for invalid data, too (the above would raise an exception instantly if there is a YAML syntax error or if the data doesn't match the schema).VSTS has a similar online text editor for editing and previewing markdown files.
I have a web hook that builds the markdown using mkdocs so we can have a nice web interface for others who do not have edit privileges.
VSTS recently released a wiki feature so I’m considering moving to that instead of mkdocs.
I think people see the value in wikis but it’s hard to sell git to someone who hasn’t used it before.
I’m always interested in this type of workflow but the problem that no one has seemed to solve is the local editing for a draft document for someone who understands markdown but not git.
.bat or .sh files? :)
A decade ago I was transitioning a small dev shop to Subversion, and daily user #2 was an non-tech accountant who was a few years away from retirement...
I didn't teach her branching. I didn't teach her merging. I didn't teach her what a trunk was, or how slicing works. She got two batch files: "Get_Newest" and "Share_Work", and presto-chango she was contributing user documentation and keeping our help-files in order inside our development project.
These days, UI competence willing, I'd just install VS Code or Atom with a markdown plugin and point it at the right GitHub repo. With smart branching and repo division I think it's a low-barrier way to get people committing via Git.
I would like to have something like NetlifyCMS but locally stored. https://www.netlifycms.org/
Now get off my lawn!
Perhaps it's less stressful and more natural for engineers to view deployment as part of "productivity." The assumption in these comments and many others I've seen across the industry is that engineers tend to view any work that requires context-switching away from coding as impeding their productivity. In startups you have to wear many hats even if you are an engineer, and it seems to me that maintaining the view that deployment is a drag on developer productivity is going to lead to a lot of unnecessary frustration. Why not just accept deployment as a part of "their work"?
>While this is a much quicker solution in the short term, an engineer will have to context switch out of their work, watch the release go out and make sure nothing goes wrong—that all hurts productivity.
Do you think this has caused a productivity increase? I would guess the additional overhead (unit tests for data files, training sessions for git, engineers doing code reviews, documentation, etc.) would cause more of a time suck.
I do like the idea of letting the stakeholder take care of it themselves on their schedule though without having to code up a GUI. The initial time investment for a GUI usually isn't that bad for us but the added maintenance is a huge PITA.
Take for example a scoreboard with targets on it. The simplest version is to have a dictionary of datetime to target. We would start with that, show the metric owner how to edit it, and let them maintain their targets. Won't need many tests, total time probably 30 mins max to get that solution set up (not including GitHub training, but that's amortized over multiple uses).
The alternative is building a data model, migration, editing interface which is probably at least 2 screens (list, edit) and 3 endpoints (list, modify, delete), needs validation on the input, needs to be put in a menu somewhere, might need a date input library on the frontend, etc.
This solution would probably end up being at least several hours all-in (code review, etc), and this is a simple data model that can easily be put in the database, there are others, with logic for example, that would need a lot more thought to get into the database. For an interface that might be used quarterly for 10 mins, it's not worth it.
In my current org I would still lean on engineers making the config changes, but if we had a group of more tech-savvy stakeholders I could see this being useful. I hope you all post an update down the road on this process.
We’re also hiring! https://www.thread.com/jobs
If you work with people who pride themselves being technology-challenged, you're doomed from the start.
If they don't fear computers and are open to new solutions - it's a piece of cake.
Github private repos for business and Github Enterprise are the same price: https://github.com/pricing
> Very sensible dudes
We also care deeply about diversity and are trying to improve our gender balance, so I'd strongly encourage any women or people of other groups traditionally under-represented in tech who might be reading this to apply, even if you don't feel you meet the requirements for the job.
It would also solve a lot of problems with mono-repos.
We just solve this with code review.
Monolithic version control systems tend to do this very well.
In my head a distributed system like GIT really needs awesome and integrated sub-module features to make it go around. Last I tried this was in-theory easy but in-practice hard with Git.
As-is I believe separate repos or custom commit scripts/hooks might get you farther :/
The point is that they are code, or a file directly loaded in, stored in source control, rather than dynamic data stored in a database that needs an editing interface.
Oh it's just other fully functional adult humans who speak the same language.
So I'm guessing the "how" is "the same way you teach any other person to use github"?