Abstract
We introduce a new metric, termed sequence-subset distance, and a new family of codes with respect to this new metric, termed sequence-subset codes, on the power set of the set of all sequences of the same length over a finite alphabet. This new metric and family of codes are motivated by the problem of designing error correcting codes for DNA-based data storage, which can be mathematically modelled as a communication channel whose inputs and outputs are sets of unordered sequences. We derive some upper bounds on the size of the sequence-subset codes including a tight bound for a special case and a Singletonlike bound, and present some constructions of such codes.