synapse-old

Commit Graph

Author	SHA1	Message	Date
Erik Johnston	51055c8c44	Allow ReplicationRestResource to be added to workers (#7515 ) This allows workers to talk to each other over HTTP replication.	2020-05-18 12:24:48 +01:00
Richard van der Hoff	4d1afb1dfe	Merge pull request #7519 from matrix-org/rav/kill_py2_code Kill off some old python 2 code	2020-05-18 10:45:30 +01:00
Richard van der Hoff	91f51c611c	remove redundant `__func__` this is a no-op under python 3	2020-05-15 19:37:41 +01:00
Richard van der Hoff	6c1f7c722f	Fix limit logic for AccountDataStream (#7384 ) Make sure that the AccountDataStream presents complete updates, in the right order. This is much the same fix as #7337 and #7358, but applied to a different stream.	2020-05-15 19:03:25 +01:00
Erik Johnston	1f36ff69e8	Move event stream handling out of slave store. (#7491 ) This allows us to have the logic on both master and workers, which is necessary to move event persistence off master. We also combine the instantiation of ID generators from DataStore and slave stores to the base worker stores. This allows us to select which process writes events independently of the master/worker splits.	2020-05-15 16:43:59 +01:00
Erik Johnston	4734a7bbe4	Move EventStream handling into default ReplicationDataHandler (#7493 ) This is so that the logic can happen on both master and workers when we move event persistence out.	2020-05-14 14:01:39 +01:00
Erik Johnston	1de36407d1	Add `instance_map` config and route replication calls (#7495 )	2020-05-14 14:00:58 +01:00
Erik Johnston	7ee24c5674	Have all instances correctly respond to REPLICATE command. (#7475 ) Before all streams were only written to from master, so only master needed to respond to `REPLICATE` commands. Before all instances wrote to the cache invalidation stream, but didn't respond to `REPLICATE`. This was a bug, which could lead to missed rows from cache invalidation stream if an instance is restarted, however all the caches would be empty in that case so it wasn't a problem.	2020-05-13 10:27:02 +01:00
Erik Johnston	8ca79613e6	Fix Redis reconnection logic (#7482 ) Proactively send out `POSITION` commands (as if we had just received a `REPLICATE`) when we connect to Redis. This is important as other instances won't notice we've connected to issue a `REPLICATE` command (unlike for direct TCP connections). This is only currently an issue if master process reconnects without restarting (if it restarts then it won't have written anything and so other instances probably won't have missed anything).	2020-05-13 09:57:15 +01:00
Amber Brown	7cb8b4bc67	Allow configuration of Synapse's cache without using synctl or environment variables (#6391 )	2020-05-11 18:45:23 +01:00
Andrew Morgan	5cf758cdd6	Merge branch 'release-v1.13.0' into develop * release-v1.13.0: Don't UPGRADE database rows RST indenting Put rollback instructions in upgrade notes Fix changelog typo Oh yeah, RST Absolute URL it is then Fix upgrade notes link Provide summary of upgrade issues in changelog. Fix ) Move next version notes from changelog to upgrade notes Changelog fixes 1.13.0rc1 Documentation on setting up redis (#7446) Rework UI Auth session validation for registration (#7455) Fix errors from malformed log line (#7454) Drop support for redis.dbid (#7450)	2020-05-11 16:46:33 +01:00
Richard van der Hoff	aa5aa6f96a	Fix errors from malformed log line (#7454 )	2020-05-07 19:51:38 +01:00
Richard van der Hoff	da9b2db3af	Drop support for redis.dbid (#7450 ) Since we only use pubsub, the dbid is irrelevant.	2020-05-07 16:46:15 +01:00
Erik Johnston	d7983b63a6	Support any process writing to cache invalidation stream. (#7436 )	2020-05-07 13:51:08 +01:00
Richard van der Hoff	62ee862119	Merge branch 'release-v1.13.0' into develop	2020-05-06 15:56:03 +01:00
Richard van der Hoff	2e0c46ca07	Merge branch 'release-v1.13.0' into develop	2020-05-06 11:58:31 +01:00
Richard van der Hoff	a8c17da245	Merge branch 'release-v1.13.0' into rav/fix_dropped_messages	2020-05-05 23:01:12 +01:00
Richard van der Hoff	1242267316	Merge branch 'release-v1.13.0' into rav/fix_dropped_messages	2020-05-05 22:38:44 +01:00
Richard van der Hoff	7f7eedbebb	Wait for a POSITION on the right connection before accepting RDATA ... otherwise we can believe we're up to date when we're not.	2020-05-05 22:38:16 +01:00
Brendan Abolivier	5b8023dc7f	Move logs about discarded RDATA to debug (#7421 )	2020-05-05 21:07:33 +02:00
Richard van der Hoff	d78265af0c	Wait to subscribe before sending REPLICATE	2020-05-05 19:31:37 +01:00
Richard van der Hoff	d5aa7d93ed	Fix catchup-on-reconnect for the Federation Stream (#7374 ) looks like we managed to break this during the refactorathon.	2020-05-05 14:15:57 +01:00
Erik Johnston	350421e058	Fix redis password support. (#7401 ) We forgot to set the password on the subscriber connection, as well as not calling super methods for overridden connectionMade/connectionLost functions.	2020-05-04 14:04:09 +01:00
Erik Johnston	0e719f2398	Thread through instance name to replication client. (#7369 ) For in memory streams when fetching updates on workers we need to query the source of the stream, which currently is hard coded to be master. This PR threads through the source instance we received via `POSITION` through to the update function in each stream, which can then be passed to the replication client for in memory streams.	2020-05-01 17:19:56 +01:00
Erik Johnston	3085cde577	Use `stream.current_token()` and remove `stream_positions()` (#7172 ) We move the processing of typing and federation replication traffic into their handlers so that `Stream.current_token()` points to a valid token. This allows us to remove `get_streams_to_replicate()` and `stream_positions()`.	2020-05-01 15:21:35 +01:00
Richard van der Hoff	b2dba06079	Workaround for assertion errors from db_query_to_update_function (#7378 ) Hopefully this is no worse than what we have on master...	2020-05-01 09:25:16 +01:00
Erik Johnston	37f6823f5b	Add instance name to RDATA/POSITION commands (#7364 ) This is primarily for allowing us to send those commands from workers, but for now simply allows us to ignore echoed RDATA/POSITION commands that we sent (we get echoes of sent commands when using redis). Currently we log a WARNING on the master process every time we receive an echoed RDATA.	2020-04-29 16:23:08 +01:00
Erik Johnston	3eab76ad43	Don't relay REMOTE_SERVER_UP cmds to same conn. (#7352 ) For direct TCP connections we need the master to relay REMOTE_SERVER_UP commands to the other connections so that all instances get notified about it. The old implementation just relayed to all connections, assuming that sending back to the original sender of the command was safe. This is not true for redis, where commands sent get echoed back to the sender, which was causing master to effectively infinite loop sending and then re-receiving REMOTE_SERVER_UP commands that it sent. The fix is to ensure that we only relay to other connections and not to the connection we received the notification from. Fixes #7334.	2020-04-29 14:10:59 +01:00
Richard van der Hoff	c2e1a2110f	Fix limit logic for EventsStream (#7358 ) * Factor out functions for injecting events into database I want to add some more flexibility to the tools for injecting events into the database, and I don't want to clutter up HomeserverTestCase with them, so let's factor them out to a new file. * Rework TestReplicationDataHandler This wasn't very easy to work with: the mock wrapping was largely superfluous, and it's useful to be able to inspect the received rows, and clear out the received list. * Fix AssertionErrors being thrown by EventsStream Part of the problem was that there was an off-by-one error in the assertion, but also the limit logic was too simple. Fix it all up and add some tests.	2020-04-29 12:30:36 +01:00
Erik Johnston	38919b521e	Run replication streamers on workers (#7146 ) Currently we never write to streams from workers, but that will change soon	2020-04-28 13:34:12 +01:00
Richard van der Hoff	ce428a1abe	Fix EventsStream raising assertions when it falls behind Figuring out how to correctly limit updates from this stream without dropping entries is far more complicated than just counting the number of rows being returned. We need to consider each query separately and, if any one query hits the limit, truncate the results from the others. I think this also fixes some potentially long-standing bugs where events or state changes could get missed if we hit the limit on either query.	2020-04-24 13:59:21 +01:00
Richard van der Hoff	9cbdfb3a2f	Make it clear that the limit for an update_function is a target	2020-04-23 15:45:12 +01:00
Richard van der Hoff	23b28266ac	Remove 'limit' param from `get_repl_stream_updates` API there doesn't seem to be much point in passing this limit all around, since both sides agree it's meant to be 100.	2020-04-23 15:44:35 +01:00
Richard van der Hoff	71a1abb8a1	Stop the master relaying USER_SYNC for other workers (#7318 ) Long story short: if we're handling presence on the current worker, we shouldn't be sending USER_SYNC commands over replication. In an attempt to figure out what is going on here, I ended up refactoring some bits of the presencehandler code, so the first 4 commits here are non-functional refactors to move this code slightly closer to sanity. (There's still plenty to do here :/). Suggest reviewing individual commits. Fixes (I hope) #7257.	2020-04-22 22:39:04 +01:00
Erik Johnston	841c581c40	Fix replication metrics when using redis (#7325 )	2020-04-22 16:26:19 +01:00
Richard van der Hoff	82d8b1dd1f	Another go at fixing one-word commands (#7326 ) I messed this up last time I tried (#7239 / `e13c6c7`).	2020-04-22 14:34:31 +01:00
Erik Johnston	51f7eaf908	Add ability to run replication protocol over redis. (#7040 ) This is configured via the `redis` config options.	2020-04-22 13:07:41 +01:00
Richard van der Hoff	0f8f02bc39	On catchup, process each row with its own stream id (#7286 ) Other parts of the code (such as the StreamChangeCache) assume that there will not be multiple changes with the same stream id. This code was introduced in #7024, and I hope this fixes #7206.	2020-04-20 11:43:29 +01:00
Richard van der Hoff	67ff7b8ba0	Improve type checking in `replication.tcp.Stream` (#7291 ) The general idea here is to get rid of the type: ignore annotations on all of the current_token and update_function assignments, which would have caught #7290. After a bit of experimentation, it seems like the least-awful way to do this is to pass the offending functions in as parameters to the Stream constructor. Unfortunately that means that the concrete implementations no longer have the same constructor signature as Stream itself, which means that it gets hard to correctly annotate STREAMS_MAP. I've also introduced a couple of new types, to take out some duplication.	2020-04-17 14:49:55 +01:00
Richard van der Hoff	d7d42387f5	Fix 'generator object is not subscriptable' error (#7290 ) Some of the query functions return generators rather than lists, so we can't index into the result. Happily we already have a copy of the results. (think this was introduced in #7024)	2020-04-16 14:37:06 +01:00
Richard van der Hoff	e13c6c7a96	Handle one-word replication commands correctly `REPLICATE` is now a valid command, and it's nice if you can issue it from the console without remembering to call it `REPLICATE ` with a trailing space.	2020-04-07 17:43:46 +01:00
Richard van der Hoff	c3e4b4edb2	Fix warnings about not calling superclass constructor Separate `SimpleCommand` from `Command`, so that things which don't want to use the `data` property don't have to, and thus fix the warnings PyCharm was giving me about not calling `__init__` in the base class.	2020-04-07 17:40:22 +01:00
Richard van der Hoff	6a519a0ca0	Remove vestigal references to SYNC replication command We've ripped pretty much all of this out: let's remove the remains.	2020-04-07 17:40:07 +01:00
Erik Johnston	ce72355d7f	Fix race in replication (#7226 ) Fixes a race between handling `POSITION` and `RDATA` commands. We do this by simply linearizing handling of them.	2020-04-07 11:01:04 +01:00
Erik Johnston	82498ee901	Move server command handling out of TCP protocol (#7187 ) This completes the merging of server and client command processing.	2020-04-07 10:51:07 +01:00
Erik Johnston	5016b162fc	Move client command handling out of TCP protocol (#7185 ) The aim here is to move the command handling out of the TCP protocol classes and to also merge the client and server command handling (so that we can reuse them for redis protocol). This PR simply moves the client paths to the new `ReplicationCommandHandler`, a future PR will move the server paths too.	2020-04-06 09:58:42 +01:00
Erik Johnston	dfa0782254	Remove connections per replication stream metric. (#7195 ) This broke in a recent PR (#7024) and is no longer useful due to all replication clients implicitly subscribing to all streams, so let's just remove it.	2020-04-01 10:40:46 +01:00
Erik Johnston	4f21c33be3	Remove usage of "conn_id" for presence. (#7128 ) * Remove `conn_id` usage for UserSyncCommand. Each tcp replication connection is assigned a "conn_id", which is used to give an ID to a remotely connected worker. In a redis world, there will no longer be a one to one mapping between connection and instance, so instead we need to replace such usages with an ID generated by the remote instances and included in the replicaiton commands. This really only effects UserSyncCommand. * Add CLEAR_USER_SYNCS command that is sent on shutdown. This should help with the case where a synchrotron gets restarted gracefully, rather than rely on 5 minute timeout.	2020-03-30 16:37:24 +01:00
Erik Johnston	4cff617df1	Move catchup of replication streams to worker. (#7024 ) This changes the replication protocol so that the server does not send down `RDATA` for rows that happened before the client connected. Instead, the server will send a `POSITION` and clients then query the database (or master out of band) to get up to date.	2020-03-25 14:54:01 +00:00
Richard van der Hoff	a564b92d37	Convert `*StreamRow` classes to inner classes (#7116 ) This just helps keep the rows closer to their streams, so that it's easier to see what the format of each stream is.	2020-03-23 13:59:11 +00:00
Richard van der Hoff	b3cee0ce67	Fix processing of `groups` stream, and use symbolic names for streams (#7117 ) `groups` != `receipts` Introduced in #6964	2020-03-23 11:39:36 +00:00
Erik Johnston	fdb1344716	Remove concept of a non-limited stream. (#7011 )	2020-03-20 14:40:47 +00:00
Erik Johnston	a319cb1dd1	Change device list streams to have one row per ID (#7010 ) * Add 'device_lists_outbound_pokes' as extra table. This makes sure we check all the relevant tables to get the current max stream ID. Currently not doing so isn't problematic as the max stream ID in `device_lists_outbound_pokes` is the same as in `device_lists_stream`, however that will change. * Change device lists stream to have one row per id. This will make it possible to process the streams more incrementally, avoiding having to process large chunks at once. * Change device list replication to match new semantics. Instead of sending down batches of user ID/host tuples, send down a row per entity (user ID or host). * Newsfile * Remove handling of multiple rows per ID * Fix worker handling * Comments from review	2020-03-19 11:36:53 +00:00
Erik Johnston	6e6476ef07	Comments from review	2020-03-18 10:13:55 +00:00
Richard van der Hoff	78a15b1f9d	Store room_versions in EventBase objects (#6875 ) This is a bit fiddly because it all has to be done on one fell swoop: * Wherever we create a new event, pass in the room version (and check it matches the format version) * When we prune an event, use the room version of the unpruned event to create the pruned version. * When we pass an event over the replication protocol, pass the room version over alongside it, and use it when deserialising the event again.	2020-03-05 15:46:44 +00:00
Erik Johnston	9ce4e344a8	Change device list replication to match new semantics. Instead of sending down batches of user ID/host tuples, send down a row per entity (user ID or host).	2020-02-28 11:25:34 +00:00
Erik Johnston	c3c6c0e622	Add 'device_lists_outbound_pokes' as extra table. This makes sure we check all the relevant tables to get the current max stream ID. Currently not doing so isn't problematic as the max stream ID in `device_lists_outbound_pokes` is the same as in `device_lists_stream`, however that will change.	2020-02-28 11:15:11 +00:00
Richard van der Hoff	3e99528f2b	Store room version on invite (#6983 ) When we get an invite over federation, store the room version in the rooms table. The general idea here is that, when we pull the invite out again, we'll want to know what room_version it belongs to (so that we can later redact it if need be). So we need to store it somewhere...	2020-02-26 16:58:33 +00:00
Erik Johnston	1f773eec91	Port PresenceHandler to async/await (#6991 )	2020-02-26 15:33:26 +00:00
Erik Johnston	bbf8886a05	Merge worker apps into one. (#6964 )	2020-02-25 16:56:55 +00:00
Erik Johnston	0bd8cf435e	Increase MAX_EVENTS_BEHIND for replication clients	2020-02-21 09:04:33 +00:00
Erik Johnston	de2d267375	Allow moving group read APIs to workers (#6866 )	2020-02-07 11:14:19 +00:00
Erik Johnston	c3d4ad8afd	Fix sending server up commands from workers (#6811 ) Co-authored-by: Andrew Morgan <1342360+anoadragon453@users.noreply.github.com>	2020-01-30 16:42:11 +00:00
Erik Johnston	e17a110661	Detect unknown remote devices and mark cache as stale (#6776 ) We just mark the fact that the cache may be stale in the database for now.	2020-01-28 14:43:21 +00:00
Erik Johnston	d5275fc55f	Propagate cache invalidates from workers to other workers. (#6748 ) Currently if a worker invalidates a cache it will be streamed to master, which then didn't forward those to other workers.	2020-01-27 13:47:50 +00:00
Erik Johnston	5d7a6ad223	Allow streaming cache invalidate all to workers. (#6749 )	2020-01-22 10:37:00 +00:00
Erik Johnston	a8a50f5b57	Wake up transaction queue when remote server comes back online (#6706 ) This will be used to retry outbound transactions to a remote server if we think it might have come back up.	2020-01-17 10:27:19 +00:00
Erik Johnston	48c3a96886	Port synapse.replication.tcp to async/await (#6666 ) * Port synapse.replication.tcp to async/await * Newsfile * Correctly document type of on_<FOO> functions as async * Don't be overenthusiastic with the asyncing....	2020-01-16 09:16:12 +00:00
Erik Johnston	28c98e51ff	Add `local_current_membership` table (#6655 ) Currently we rely on `current_state_events` to figure out what rooms a user was in and their last membership event in there. However, if the server leaves the room then the table may be cleaned up and that information is lost. So lets add a table that separately holds that information.	2020-01-15 14:59:33 +00:00
Erik Johnston	e8b68a4e4b	Fixup synapse.replication to pass mypy checks (#6667 )	2020-01-14 14:08:06 +00:00
Richard van der Hoff	6964ea095b	Reduce the reconnect time when replication fails. (#6617 )	2020-01-03 14:19:09 +00:00
Erik Johnston	fa780e9721	Change EventContext to use the Storage class (#6564 )	2019-12-20 10:32:02 +00:00
Erik Johnston	9a4fb457cf	Change DataStores to accept 'database' param.	2019-12-06 13:30:06 +00:00
Erik Johnston	a7f20500ff	_CURRENT_STATE_CACHE_NAME is public	2019-12-04 15:45:42 +00:00
Erik Johnston	1056d6885a	Move cache invalidation to main data store	2019-12-04 15:21:14 +00:00
Erik Johnston	2173785f0d	Propagate reason in remotely rejected invites	2019-11-28 11:31:56 +00:00
Andrew Morgan	a8175d0f96	Prevent account_data content from being sent over TCP replication (#6333 )	2019-11-26 13:58:39 +00:00
Erik Johnston	f9f1c8acbb	Merge pull request #6332 from matrix-org/erikj/query_devices_fix Fix caching devices for remote servers in worker.	2019-11-26 12:56:05 +00:00
Erik Johnston	35f9165e96	Fixup docs	2019-11-26 12:04:48 +00:00
Andrew Morgan	cd96b4586f	lint	2019-11-08 15:45:45 +00:00
Andrew Morgan	c4bdf2d785	Remove content from being sent for account data rdata stream	2019-11-08 15:44:02 +00:00
Andrew Morgan	1fe3cc2c9c	Address review comments	2019-11-06 14:54:24 +00:00
Andrew Morgan	4059d61e26	Don't forget to ratelimit calls outside of RegistrationHandler	2019-11-06 12:01:54 +00:00
Erik Johnston	c16e192e2f	Fix caching devices for remote servers in worker. When the `/keys/query` API is hit on client_reader worker Synapse may decide that it needs to resync some remote deivces. Usually this happens on master, and then gets cached. However, that fails on workers and so it falls back to fetching devices from remotes directly, which may in turn fail if the remote is down.	2019-11-05 15:49:43 +00:00
Richard van der Hoff	cc6243b4c0	document the REPLICATE command a bit better (#6305 ) since I found myself wonder how it works	2019-11-04 12:40:18 +00:00
Hubert Chathi	9c94b48bf1	Merge branch 'develop' into uhoreg/cross_signing_fix_workers_notify	2019-10-31 12:32:07 -04:00
Hubert Chathi	f7e4a582ef	clean up code a bit	2019-10-31 12:01:00 -04:00
Andrew Morgan	54fef094b3	Remove usage of deprecated logger.warn method from codebase (#6271 ) Replace every instance of `logger.warn` with `logger.warning` as the former is deprecated.	2019-10-31 10:23:24 +00:00
Hubert Chathi	998f7fe7d4	make user signatures a separate stream	2019-10-30 17:22:52 -04:00
Hubert Chathi	670972c0e1	Merge branch 'develop' into uhoreg/cross_signing_fix_workers_notify	2019-10-30 16:46:31 -04:00
Erik Johnston	e577a4b2ad	Port replication http server endpoints to async/await	2019-10-29 13:00:51 +00:00
Hubert Chathi	8ac766c44a	make notification of signatures work with workers	2019-10-24 22:14:58 -04:00
Erik Johnston	bb6264be0b	Merge branch 'develop' of github.com:matrix-org/synapse into erikj/refactor_stores	2019-10-22 10:41:18 +01:00
Erik Johnston	c66a06ac6b	Move storage classes into a main "data store". This is in preparation for having multiple data stores that offer different functionality, e.g. splitting out state or event storage.	2019-10-21 16:05:06 +01:00
Hubert Chathi	8e86f5b65c	Merge branch 'develop' into uhoreg/e2e_cross-signing_merged	2019-09-07 13:20:34 -04:00
Jorik Schellekens	f7c873a643	Trace how long it takes for the send trasaction to complete, including retrys (#5986 )	2019-09-05 17:44:55 +01:00
Jorik Schellekens	909827b422	Add opentracing to all client servlets (#5983 )	2019-09-05 14:46:04 +01:00
Hubert Chathi	a22d58c96c	add user signature stream change cache to slaved device store	2019-09-04 19:32:35 -04:00
Andrew Morgan	b736c6cd3a	Remove bind_email and bind_msisdn (#5964 ) Removes the `bind_email` and `bind_msisdn` parameters from the `/register` C/S API endpoint as per [MSC2140: Terms of Service for ISes and IMs](https://github.com/matrix-org/matrix-doc/pull/2140/files#diff-c03a26de5ac40fb532de19cb7fc2aaf7R107).	2019-09-04 18:24:23 +01:00
Andrew Morgan	4548d1f87e	Remove unnecessary parentheses around return statements (#5931 ) Python will return a tuple whether there are parentheses around the returned values or not. I'm just sick of my editor complaining about this all over the place :)	2019-08-30 16:28:26 +01:00

1 2 3 4 5 ...

482 Commits