Showing posts with label dynamodb. Show all posts
Showing posts with label dynamodb. Show all posts

Friday, May 18, 2018

Why we have a memcached-dynamo proxy

Previously, I mentioned that we have a proxy that speaks the memcached protocol to clients and stores data in dynamodb, “for reasons.”  Today, it’s time to talk about the architecture, and those reasons.

Friday, December 12, 2014

DynamoDB's Greatest Missing Feature

I wrote a while ago about DynamoDB in the trenches, but that was in 2012 and now 2015 is nigh.

The real pain point of DynamoDB is that everything is a hash key.  It's not sorted, at all.  Ever.

A local secondary index isn't the answer; it's a generalization of range keys, and as such, is still subordinate to the hash key of the table.  An LSI does not cover items with different hash keys.

A global secondary index isn't helpful in this regard, either.  It is essentially building and maintaining a second table for you, where its hash key is the attribute the index is on, and then it stores the key of the target element (and optionally, any other attributes from that element) from the 'parent' table.  But as a hash key, it doesn't support ordering...  The only thing that can be done with the index is a Scan operation.

DynamoDB still offers nothing else.

Some ideas that would seem, on their surface, to be perfect for DynamoDB actually don't work out well, because there's no efficient query for ordered metadata.  Most frequently, this bites me for data that expires, like sessions.  Sessions!  The use case that's so blindingly obvious, the PHP SDK includes a DynamoDB session handler.

Update (2023-01-28): DynamoDB has support for an 'expiration time' attribute these days.  We are using it for sessions now.  Due to encoding issues, our sessions are stored as binary instead of string type, but otherwise, it's compatible with the PHP SDK's data format.  Our programming languages actually use memcached but the data is stored in DynamoDB. End of update.

We store SES message data in there to correlate bounces with the responsible party.  We have a nightly job that expires the old junk, and just live with it sucking up a ton of read capacity when it has to do that.  On the bright side, unlike sessions, the website won't grind to a halt (or log people out erroneously) if the table gets wedged.

I'd like to put some rarely-changing configuration data in there (it's in SimpleDB and cached in memcache because SimpleDB had erratic query performance), but I wouldn't be able to efficiently search over it when I wanted to look up an entry for editing in the admin interface.  And really, I still want to put that in LDAP or something.

Honestly, if all I get with DynamoDB is a blob of stuff and a hash key, why not just store it as a file in S3?  How is DynamoDB even a database if it doesn't actually do anything with the data, and doesn't let you have any metadata?  Am I the crazy one, or is that the NoSQL crowd?

Related: Tim Gross on Falling In And Out Of Love with DynamoDB.

Tuesday, November 27, 2012

DynamoDB in the Trenches

Amazon's DynamoDB, as they're happy to tell you, is an SSD-backed NoSQL storage service with provisioned throughput.  Data consists of three basic types (number, string, and binary) in either scalar or set forms.  (A set contains any number of unique values of its parent type, in no particular order.)  All lookups are done by hash keys, optionally with a range as sub-key; the former effectively defines a root type, and the latter is something like a 1:Many relation.  The hash key is the foreign key to the parent object, and the range key defines an ordering of the collection.

But you knew all that; there are two additional points I want to add to the documentation.

1. Update is Upsert

DynamoDB's update operation actually behaves as upsert—if you update a nonexistent item, the attributes you updated will be created as the only attributes on the item.  If this would result in an invalid item as a whole, then you want to use the expected-value mechanism to make sure the item is really there before the update applies.

2. No Attribute Indexing or Querying

NoSQL is at once the greatest draw to non-relational storage, and also its biggest drawback.  On DynamoDB, there's no query language.  You can get items by key cheaply, or scan the whole table expensively, and there's nothing in between.  There's no API for indexing any attributes, so there's no API for querying by index, either.  You can't even query on the range of a range key independently of the hash (e.g. "find all the posts today, regardless of topic" on a table keyed by a topic-hash and date-range.)

If you need lookup by attribute-equals more than you need consistent query performance, then you can use SimpleDB.  RDS could be a decent option, especially if you want ordered lookups (as in DELETE FROM sessions WHERE expires < NOW();—when the primary key is "id".)

A not-so-good option would be to add another DynamoDB table keyed by attribute-values and containing sets of your main table's hash keys—but you can't update multiple DynamoDB tables transactionally, so this is more prone to corruption than other methods.

And if you want to pull together two sources of asynchronous events, and act on the result once both events have occurred, then the Simple Workflow service might work.  (I realized this about 98% of the way through a different solution, so I stuck with that.  I might have been able to store the half-complete event into the workflow state instead, no DynamoDB needed, but since I didn't walk that path, I can't vouch for it.)

Monday, July 30, 2012

DynamoDB Performance

Things I learned this past week:
  1. AWS Signature V4 no longer requires temporary credentials.  If you aren’t caching your tokens, this can give you a nice speedup because it cuts IAM/STS out of the connection sequence.
  2. AWS service endpoints are SSL.  If you make a lot of fresh connections, you may pay a lot of overhead per connection.
  3. Net::Amazon::DynamoDB and CGI are terrible things to mix.
Read on for details.