I would like to ask everyone something: what is the format of the parameters passed when calling the ‘/api/storage’ endpoint, and what is the relationship between ‘storage_id’ and ‘vector encoding’, and between ‘Image encoding’ and ‘identity_id’?
'storage_id' and 'vector encoding', 'Image encoding' and 'identity_id'
Archived from the BrainFrame forum. Posted Jun 16, 2021. 1 reply · 457 views Translated into English from the original Chinese posts. Read the original
What is the relationship between storage_id and vector_encoding, image_encoding and identity_id
Okay, that’s a good question. Think of Identity as a specific person. That person can have many things associated with it. You might want to associate a ‘face’ and a ‘person’ and a ‘license_plate’ with the same Identity. Because of that, we allow creation of identities, and then the association of an identity with a specific image.
These “images” are templates. For example, think about a face recognition application. It needs the user to upload example images of “Alex”'s face in order to recognize “Alex”. There might even be multiple template images- so the relationship between # of template images and an Identity is NOT 1:1.
When you upload face images to brainframe (documented here Identities) what happens is BrainFrame will try to extract a “encoding” that describes that face. That “encoding” is a vector, produced by a machine learning model inside of a capsule. BrainFrame can use that vector/encoding to judge the similarity of faces in a video stream to known faces from a template.
HOWEVER- sometimes, there are objects that can be “recognized” that are NOT faces. For example, in the case of a fiducial tag (example image: https://i.imgur.com/UMoMbBa.png), these tags can automatically be converted to a number or a vector. For those cases, BrainFrame allows you to directly upload a “vector” to associate with an Identity, that way, if a capsule is able to detect that Vector, BrainFrame will then associate it with an identity and return it through the API.
To summarize:
Identities can have many ‘templates’ associated with them. Different pictures of the same persons face, maybe pictures of a license plate, etc.
BrainFrame recognizes things by converting images into vectors, using whatever method the capsule outputs. It then compares vectors to see how similar they are, and matches similar vectors with the known identity in the database.
Imagine how you recognize someone's identity in real life. Say you have a friend called Zhang San. How do you decide whether the person in front of you is Zhang San? You search your memory for Zhang San's features and compare them with the person in front of you to decide whether it is the same person. The features might be his face, his clothes, his height and weight, and so on. You may also hold more than one memory of Zhang San — Zhang San with different hairstyles, at different ages, in different clothes — and you compare these different remembered versions with the person in front of you to reach your judgment.
Unlike real life, in artificial intelligence — taking face recognition as the example — the usual method is to compute a vector for a face image with an algorithm, and then compute the difference between that vector and the target face's vector to decide whether the two faces belong to the same person. Different algorithms compute this differently, of course. For example, the QR code in that picture corresponds not to a vector but simply to a number, and the algorithm uses the number to decide whether it is the same QR code.
In BrainFrame, an Identity represents an identity — it can be a person, a vehicle or a license plate — and every Identity has a corresponding identity_id. In the real-life example above, the identity_id is Zhang San. image_encoding and vector_encoding are the ways of describing or recognizing an Identity; in the example above they could be Zhang San's face, clothes, height and weight, and so on. So the relationship between identity_id and image_encoding / vector_encoding is not 1:1: one identity_id may correspond to several image_encodings or vector_encodings.
image_encoding and vector_encoding are different. For an image_encoding, BrainFrame calls a different capsule depending on its type (for example face or license plate) to compute a vector from the picture, so an image_encoding must correspond to a picture. All pictures in BrainFrame are stored in the database and each has a corresponding storage_id, so to create a new image_encoding you first upload the picture to BrainFrame's database to get a storage_id, and then use it to create the image_encoding. A vector_encoding is the vector itself: you can upload the vector directly instead of a picture, so it can be created without a storage_id. But you need to make sure that the algorithm that generated the vector is the same one BrainFrame uses, because even for the same face, different algorithms may compute different vectors.
To sum up, one identity can correspond to several image_encodings and vector_encodings, and different image_encodings and vector_encodings may correspond to different face photos or vectors of the same person, or different photos or vectors of the same license plate. When an object is detected, BrainFrame calls a different algorithm capsule depending on the type of the object to compute a vector, and compares it with the image_encodings or vector_encodings in the database to determine the object's identity.