- what happens if your processing logic changes 6 months down the line?
- for example today you decide you want to remove all extra spaces and lowercase all titles before storing them
- 6 months down the line you want to revert, what now?
- what happens if your processing logic changes 6 months down the line?
- for example today you decide you want to remove all extra spaces and lowercase all titles before storing them
- 6 months down the line you want to revert, what now?
But also store derived data. Titles, authors, dates, article texts. You need those for whatever your application does. You don't want your application logic to be working with the raw text.
> how will you handle updates to the feed?
When polling, consider using HTTP HEAD to check for changes before GET.
What you do when an article ID reappears with different content, that up to you. I think readers usually replace the old entry with the new content, silently. But it's not the only option.
- if you split the rss document into its "items" i am not sure if you can store each item separately inside postgres using that XML data type
- if you store the whole document, you end up with a problem when say the feed partially updates after a few minutes
- one suggestion i hear from r/PostgreSQL is that you store the XML blob somewhere else like hstore or something and somehow index it back to postgres. I wonder how that works conceptually / architecturally speaking
You don’t even need to store it in pg itself; stuff it into a cheaper datastore like s3 and just have the locations stored in pg.
The only thing to optimize for is cost & storage. Access/retrieval doesn’t matter for a once in 6 months process.
Also what is this revolting formatting strategy you’ve found?