Skip to content
  • Steve Wise's avatar
    RDMA/cxgb3: Support peer-2-peer connection setup · f8b0dfd1
    Steve Wise authored
    
    
    Open MPI, Intel MPI and other applications don't respect the iWARP
    requirement that the client (active) side of the connection send the
    first RDMA message.  This class of application connection setup is
    called peer-to-peer.  Typically once the connection is setup, _both_
    sides want to send data.
    
    This patch enables supporting peer-to-peer over the chelsio RNIC by
    enforcing this iWARP requirement in the driver itself as part of RDMA
    connection setup.
    
    Connection setup is extended, when the peer2peer module option is 1,
    such that the MPA initiator will send a 0B Read (the RTR) just after
    connection setup.  The MPA responder will suspend SQ processing until
    the RTR message is received and reply-to.
    
    In the longer term, this will be handled in a standardized way by
    enhancing the MPA negotiation so peers can indicate whether they
    want/need the RTR and what type of RTR (0B read, 0B write, or 0B send)
    should be sent.  This will be done by standardizing a few bits of the
    private data in order to negotiate all this.  However this patch
    enables peer-to-peer applications now and allows most of the required
    firmware and driver changes to be done and tested now.
    
    Design:
    
     - Add a module option, peer2peer, to enable this mode.
    
     - New firmware support for peer-to-peer mode:
    
    	- a new bit in the rdma_init WR to tell it to do peer-2-peer
    	  and what form of RTR message to send or expect.
    
    	- process _all_ preposted recvs before moving the connection
    	  into rdma mode.
    
    	- passive side: defer completing the rdma_init WR until all
    	  pre-posted recvs are processed.  Suspend SQ processing until
    	  the RTR is received.
    
    	- active side: expect and process the 0B read WR on offload TX
    	  queue. Defer completing the rdma_init WR until all
    	  pre-posted recvs are processed.  Suspend SQ processing until
    	  the 0B read WR is processed from the offload TX queue.
    
     - If peer2peer is set, driver posts 0B read request on offload TX
       queue just after posting the rdma_init WR to the offload TX queue.
    
     - Add CQ poll logic to ignore unsolicitied read responses.
    
    Signed-off-by: default avatarSteve Wise <swise@opengridcomputing.com>
    Signed-off-by: default avatarRoland Dreier <rolandd@cisco.com>
    f8b0dfd1