Skip to content

machinelearning:: paper link : 2007.15651.pdf (arxiv.org)

Summary

  • rather than operate on the whole image, this operates at the patch level (super useful for WSIs!!)
  • GAN with InfoNCE (contrastive) lost built in.
  • operates on unpaired data, which means there's no 1:1 coregistered images
    • in my case i have weakly paired data, where it's not 1:1, but they do correspond. But it means we can't use traditional metrics like pixel-level reconstruction.
  • style transfer seeks to make specific layers similar, but encourages multiple layers to be similar
  • use InfoNCE loss which aims to learn an embedding that associates corresponding patches to each other, while disassociating them from others

    • similar to the weakly supervised assumption that not all instances are representative

      the encoder learn to pay attention to the commonalities between two domains like object parts and shapes, and invariant to differences like textures

  • hopefully this could help us with either 1) cross-domain transfer to the testing sets (unsupervised so it's still 'legal...?') 2) learn features invariant from h&e to p53

    we find that drawing negatives internally from within the input image, rather than externally from other images in the dataset, forces the patches to better preserve the content of the input

    why not both? try internal and external negatives
    
  • in paired image-to-image translation, we can map an image from input to output domain with gans and reconstruction loss
    • here we don't necessarily need the mapping itself (THOUGH i wonder if we could do something crazy like train a cyclegan to convert the two between each other and also use it as a feature extractor)
    • paired image translation relies on cycle-consistency, whereby we assume that it's desirable to convert from input --> output --> input, but in practice this can be a restrictive heuristic
  • This work strives to maximize mutual information between associated signals
  • associated signals can be many things: an image with itself, an augmented image (traditional contrastive learning, a paired image, etc)
  • this work relies on the strong assumption that the patches with similar content will be more similar between in/output patches than patches within a single domain
    • e.g a giraffe and horse's leg are more similar to each other than the background is
    • in our case (H&E vs p53) I doubt we can make such an assumption
  • this work seeks to match specific patches between the in/out domains (find the legs in both) - this would be helpful for us to confirm that the features learned are similiar(e.g )

Ideas

Abstract

In image-to-image translation, each patch in the output should reflect the content of the corresponding patch in the input, independent of domain. We propose a straightforward method for doing so – maximizing mutual information between the two, using a framework based on contrastive learning. The method encourages two elements (corresponding patches) to map to a similar point in a learned feature space, relative to other elements (other patches) in the dataset, referred to as negatives. We explore several critical design choices for making contrastive learning effective in the image synthesis setting. Notably, we use a multilayer, patch-based approach, rather than operate on entire images. Furthermore, we draw negatives from within the input image itself, rather than from the rest of the dataset. We demonstrate that our framework enables one-sided translation in the unpaired image-to-image translation setting, while improving quality and reducing training time. In addition, our method can even be extended to the training setting where each “domain” is only a single image